Qwen2.5-Coder 32B · 10 GPUs measured first-party · llama.cpp Q4_K_M · Updated October 2026

What GPU Do You Need for Qwen2.5-Coder 32B?

Qwen2.5-Coder 32B was the local coding flagship of its generation: a dense 32B tuned for code, needing ~21GB at Q4_K_M and topping our chart at 83 tok/s on the B300. We measured it on 10 GPUs with the same pinned llama.cpp harness as every other model on this site.

Benchmarked weights: bartowski/Qwen2.5-Coder-32B-Instruct-GGUF

Fastest we measured
NVIDIA B300

NVIDIA B300

82.9 tok/s on Qwen2.5-Coder 32B, the ceiling. Measured on our bench. 288GB of VRAM, $40,000 at launch.

Pros
  • 82.9 tok/s on Qwen2.5-Coder 32B
  • 288GB, clears the Qwen2.5-Coder 32B floor
  • Rentable by the hour rather than bought
Cons
  • 1400W board rating
  • Datacenter or workstation hardware, not a retail purchase
Cheapest card that runs it
NVIDIA RTX PRO 6000 Blackwell Workstation Edition

NVIDIA RTX PRO 6000 Blackwell Workstation Edition

66.07 tok/s on Qwen2.5-Coder 32B, lowest launch price that still fits. Measured on our bench. 96GB of VRAM, $8,565 at launch.

Pros
  • 66.07 tok/s on Qwen2.5-Coder 32B
  • 96GB, clears the Qwen2.5-Coder 32B floor
  • Rentable by the hour rather than bought
Cons
  • 600W board rating
  • Datacenter or workstation hardware, not a retail purchase
82.9tok/s
Fastest: NVIDIA B300
measured
10
Cards that run Qwen2.5-Coder 32B
of 11 we have data for
1
Cards that can't run it at all
published as hard gates, not omissions
575%
Fastest vs slowest that fits
82.9 vs 12.29 tok/s

What GPU Do You Need for Qwen2.5-Coder 32B?, tok/s by GPU

NVIDIA B300
82.9 tok/s
NVIDIA B200
78.53 tok/s
NVIDIA H100 80GB HBM3
75.81 tok/s
NVIDIA H200
75.77 tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
66.07 tok/s
NVIDIA A100 80GB SXM4
43.36 tok/s
NVIDIA A100 40GB SXM4
43.14 tok/s
NVIDIA L40S
34.59 tok/s
NVIDIA A10G
21.92 tok/s
NVIDIA L4
12.29 tok/s

Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.

Efficiency: tok/s per 100W drawn

NVIDIA H100 80GB HBM3
43.39 tok/s / 100W
NVIDIA H200
29.16 tok/s / 100W
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
27.61 tok/s / 100W
NVIDIA B300
26.23 tok/s / 100W
NVIDIA A100 80GB SXM4
24.92 tok/s / 100W
NVIDIA B200
21.65 tok/s / 100W
NVIDIA A100 40GB SXM4
20.78 tok/s / 100W
NVIDIA L4
18.29 tok/s / 100W
NVIDIA A10G
16.63 tok/s / 100W
NVIDIA L40S
13.64 tok/s / 100W

Power is the average pulled during the run, sampled at 1Hz. The fastest card is often not the one here, and for anything left running this is the number that shows up on the bill.

Value: tok/s per $1,000 of MSRP

NVIDIA A10G
7.83 tok/s / $1k
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
7.71 tok/s / $1k
NVIDIA L4
4.92 tok/s / $1k
NVIDIA L40S
4.61 tok/s / $1k
NVIDIA A100 40GB SXM4
3.6 tok/s / $1k
NVIDIA A100 80GB SXM4
2.55 tok/s / $1k
NVIDIA H100 80GB HBM3
2.53 tok/s / $1k
NVIDIA H200
2.44 tok/s / $1k
NVIDIA B300
2.07 tok/s / $1k
NVIDIA B200
1.96 tok/s / $1k

Launch price, not street price, so it ages. A speed leaderboard always crowns the most expensive card; this is the counterweight.

Qwen2.5-Coder 32B. Measured generation speed by GPU

NVIDIA B30082.9
NVIDIA B20078.53
NVIDIA H100 80GB HBM375.81
NVIDIA H20075.77
NVIDIA RTX PRO 6000 Blackwell Workstation Edition66.07
NVIDIA A100 80GB SXM443.36
NVIDIA A100 40GB SXM443.14
NVIDIA L40S34.59
NVIDIA A10G21.92
NVIDIA L412.29
GPUtok/sPrompt t/stok/WAvg power
NVIDIA B30082.91336.70.26316.1 W
NVIDIA B20078.532558.40.22362.8 W
NVIDIA H100 80GB HBM375.812295.60.43174.7 W
NVIDIA H20075.772289.70.29259.8 W
NVIDIA RTX PRO 6000 Blackwell Workstation Edition66.073417.20.28239.3 W
NVIDIA A100 80GB SXM443.361188.40.25174.0 W
NVIDIA A100 40GB SXM443.141150.80.21207.6 W
NVIDIA L40S34.592378.60.14253.5 W
NVIDIA A10G21.92801.90.17131.8 W
NVIDIA L412.29702.60.1867.2 W

Where a dense 32B coder fits now. Two ways to use this model well. First, as the strong single model on a 24GB card: 24GB clears the ~21GB floor, and 30-40 tok/s-class generation is comfortable for chat-style coding. Second, and more interesting: as one voice in a consensus setup. Running the same prompt through two or three different 32B-class models, this, QwQ, something outside the Qwen family, and reconciling the answers gets you meaningfully better results than any single model. That's my current view of the best local strategy: dense 32B models as ensemble members, not soloists. Since the Qwen3 Coder MoE arrived, the MoE is the daily driver and this is the second opinion.

A value note the chart hides. The B300 tops the board at 83 tok/s, but dense 32B generation is bandwidth-bound and single-stream, so datacenter silicon is mostly wasted on it. On the sister dense model (Qwen3 32B) we measured an RTX 5090 at 71 tok/s against the B300's 84: a consumer card delivering 85% of the flagship datacenter chip's output on this workload. Even a used RTX 3090 holds 38 tok/s. Renting big iron for one-at-a-time Q4 inference is the least efficient way to spend AI money, big cards earn their rate on batch serving, not solo chat.

About Qwen2.5-Coder 32B. Qwen2.5-Coder 32B: from Qwen, 33B parameters, on Hugging Face since November 2024, Apache 2.0 licence. 1,890,015 downloads in the last 30 days and 1 community quantizations.

How it compares. H100 80GB HBM3: Qwen2.5-Coder 32B 75.81 tok/s, Qwen3-32B 74.07 (33B), DeepSeek-R1 Distill 32B 75.8 (33B), Qwen2.5-32B 73.5 (33B), Nemotron 3.5 Lightning 30B A3B 323.6 (32B). 1 of 4 beat Qwen2.5-Coder 32B here. Nemotron 3.5 Lightning 30B A3B is a mixture-of-experts, so per token it computes only a slice of its size.

Cost on a rented GPU. 1M generated tokens of Qwen2.5-Coder 32B: $3.04 on a A100 40GB SXM4 ($0.47/hr, 6.4 hours), $23.25 on a B300 ($6.94/hr, 3.4 hours, 7.7x the cost).

Qwen2.5-Coder 32B: cost per 1M generated tokens on rented GPUs

NVIDIA A100 40GB SXM4$0.47/hr
NVIDIA RTX PRO 6000 Blackwell Workstation Edition$1.08/hr
NVIDIA A100 80GB SXM4$0.95/hr
NVIDIA L40S$0.79/hr
NVIDIA H100 80GB HBM3$2.14/hr
NVIDIA L4$0.44/hr
NVIDIA H200$3.59/hr
NVIDIA B200$5.98/hr
NVIDIA B300$6.94/hr
GPUCheapest rateSpeed (tok/s)Cost per 1M generated tokens
NVIDIA A100 40GB SXM4$0.47/hr43.14$3.04
NVIDIA RTX PRO 6000 Blackwell Workstation Edition$1.08/hr66.07$4.52
NVIDIA A100 80GB SXM4$0.95/hr43.36$6.07
NVIDIA L40S$0.79/hr34.59$6.34
NVIDIA H100 80GB HBM3$2.14/hr75.81$7.83
NVIDIA L4$0.44/hr12.29$9.94
NVIDIA H200$3.59/hr75.77$13.16
NVIDIA B200$5.98/hr78.53$21.15
NVIDIA B300$6.94/hr82.9$23.25

Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.

Speed tiers for Qwen2.5-Coder 32B. 30+ tok/s: 8 (B300, B200, H100 80GB HBM3); 10-30 tok/s: 2 (A10G, L4). 30 tok/s is roughly where replies outpace reading.

Reading your prompt. Before Qwen2.5-Coder 32B writes anything it reads the input: 3417.2 tok/s on the RTX PRO 6000 Blackwell Workstation Edition (1.2s for a 4,000-token prompt), 702.6 on the L4 (5.7s). Long documents and big code files feel this number more than the generation speed.

VRAM for Qwen2.5-Coder 32B. Measured peak 19.2GB, so 24GB is the smallest common card size; smallest card it ran on: A10G (24GB). With long context: Q4_K_M 21GB (tested), Q2_K 14GB, Q3_K_M 17GB, Q5_K_M 25GB, Q6_K 29GB.

Power on Qwen2.5-Coder 32B. Most efficient: H100 80GB HBM3, 175W, 0.64 kWh per 1M generated tokens. Hungriest: B200, 363W, 1.28 kWh. At $0.15/kWh: $0.096 per 1M generated tokens.

Our verdict

Qwen2.5-Coder 32B: 83 tok/s peak, ~21GB floor, a dense coding brain that any 24GB card can hold. Our advice: run it as the strong second opinion in a consensus setup alongside the faster Qwen3 Coder MoE, and don't rent datacenter silicon for single-stream Q4 inference, a 24GB consumer card does half the B300's speed at a sliver of the cost.

FAQ

What GPU do I need for Qwen2.5-Coder 32B?
24GB minimum. Measured peak was ~21GB at Q4_K_M. RX 7900 XTX, RTX 3090/3090 Ti, RTX 4090. Nothing smaller fits it at this quant.
Is it worth renting a datacenter GPU for?
For single-user inference, no. Dense 32B generation is bandwidth-bound: on the sister Qwen3 32B, an RTX 5090 measured 71 tok/s against the B300's 84, 85% of the output from a card you can own. Datacenter cards earn their rates on batch serving, not solo chat.
Qwen2.5-Coder 32B or Qwen3 Coder 30B-A3B?
The MoE for daily use, nearly 4× the generation speed (318 vs 83 tok/s at each chart's top) in the same VRAM class. Keep the dense 32B as a second opinion on hard problems: different architecture, different failure modes.
What's a consensus setup?
Running the same prompt through several strong models, say this, QwQ 32B, and a non-Qwen 32B, and reconciling the answers. In our experience it beats any single local model, and all three fit a 24GB card one at a time.
How fast is it on a consumer card?
We measured its sister dense model (Qwen3 32B) at 71 tok/s on an RTX 5090, 44 on a 4090 and 38 on a 3090. Expect this model to land in the same bands. The 5090 makes dense 32B genuinely comfortable; older 24GB cards are chat-speed but slow for agent loops.