Qwen2.5-Coder 32B · 10 GPUs measured first-party · llama.cpp Q4_K_M · Updated October 2026
Qwen2.5-Coder 32B was the local coding flagship of its generation: a dense 32B tuned for code, needing ~21GB at Q4_K_M and topping our chart at 83 tok/s on the B300. We measured it on 10 GPUs with the same pinned llama.cpp harness as every other model on this site.
Benchmarked weights: bartowski/Qwen2.5-Coder-32B-Instruct-GGUF

82.9 tok/s on Qwen2.5-Coder 32B, the ceiling. Measured on our bench. 288GB of VRAM, $40,000 at launch.

66.07 tok/s on Qwen2.5-Coder 32B, lowest launch price that still fits. Measured on our bench. 96GB of VRAM, $8,565 at launch.
What GPU Do You Need for Qwen2.5-Coder 32B?, tok/s by GPU
Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.
Efficiency: tok/s per 100W drawn
Power is the average pulled during the run, sampled at 1Hz. The fastest card is often not the one here, and for anything left running this is the number that shows up on the bill.
Value: tok/s per $1,000 of MSRP
Launch price, not street price, so it ages. A speed leaderboard always crowns the most expensive card; this is the counterweight.
Qwen2.5-Coder 32B. Measured generation speed by GPU
| GPU | tok/s | Prompt t/s | tok/W | Avg power |
|---|---|---|---|---|
| NVIDIA B300 | 82.9 | 1336.7 | 0.26 | 316.1 W |
| NVIDIA B200 | 78.53 | 2558.4 | 0.22 | 362.8 W |
| NVIDIA H100 80GB HBM3 | 75.81 | 2295.6 | 0.43 | 174.7 W |
| NVIDIA H200 | 75.77 | 2289.7 | 0.29 | 259.8 W |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 66.07 | 3417.2 | 0.28 | 239.3 W |
| NVIDIA A100 80GB SXM4 | 43.36 | 1188.4 | 0.25 | 174.0 W |
| NVIDIA A100 40GB SXM4 | 43.14 | 1150.8 | 0.21 | 207.6 W |
| NVIDIA L40S | 34.59 | 2378.6 | 0.14 | 253.5 W |
| NVIDIA A10G | 21.92 | 801.9 | 0.17 | 131.8 W |
| NVIDIA L4 | 12.29 | 702.6 | 0.18 | 67.2 W |
Where a dense 32B coder fits now. Two ways to use this model well. First, as the strong single model on a 24GB card: 24GB clears the ~21GB floor, and 30-40 tok/s-class generation is comfortable for chat-style coding. Second, and more interesting: as one voice in a consensus setup. Running the same prompt through two or three different 32B-class models, this, QwQ, something outside the Qwen family, and reconciling the answers gets you meaningfully better results than any single model. That's my current view of the best local strategy: dense 32B models as ensemble members, not soloists. Since the Qwen3 Coder MoE arrived, the MoE is the daily driver and this is the second opinion.
A value note the chart hides. The B300 tops the board at 83 tok/s, but dense 32B generation is bandwidth-bound and single-stream, so datacenter silicon is mostly wasted on it. On the sister dense model (Qwen3 32B) we measured an RTX 5090 at 71 tok/s against the B300's 84: a consumer card delivering 85% of the flagship datacenter chip's output on this workload. Even a used RTX 3090 holds 38 tok/s. Renting big iron for one-at-a-time Q4 inference is the least efficient way to spend AI money, big cards earn their rate on batch serving, not solo chat.
About Qwen2.5-Coder 32B. Qwen2.5-Coder 32B: from Qwen, 33B parameters, on Hugging Face since November 2024, Apache 2.0 licence. 1,890,015 downloads in the last 30 days and 1 community quantizations.
How it compares. H100 80GB HBM3: Qwen2.5-Coder 32B 75.81 tok/s, Qwen3-32B 74.07 (33B), DeepSeek-R1 Distill 32B 75.8 (33B), Qwen2.5-32B 73.5 (33B), Nemotron 3.5 Lightning 30B A3B 323.6 (32B). 1 of 4 beat Qwen2.5-Coder 32B here. Nemotron 3.5 Lightning 30B A3B is a mixture-of-experts, so per token it computes only a slice of its size.
Cost on a rented GPU. 1M generated tokens of Qwen2.5-Coder 32B: $3.04 on a A100 40GB SXM4 ($0.47/hr, 6.4 hours), $23.25 on a B300 ($6.94/hr, 3.4 hours, 7.7x the cost).
Qwen2.5-Coder 32B: cost per 1M generated tokens on rented GPUs
| GPU | Cheapest rate | Speed (tok/s) | Cost per 1M generated tokens |
|---|---|---|---|
| NVIDIA A100 40GB SXM4 | $0.47/hr | 43.14 | $3.04 |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | $1.08/hr | 66.07 | $4.52 |
| NVIDIA A100 80GB SXM4 | $0.95/hr | 43.36 | $6.07 |
| NVIDIA L40S | $0.79/hr | 34.59 | $6.34 |
| NVIDIA H100 80GB HBM3 | $2.14/hr | 75.81 | $7.83 |
| NVIDIA L4 | $0.44/hr | 12.29 | $9.94 |
| NVIDIA H200 | $3.59/hr | 75.77 | $13.16 |
| NVIDIA B200 | $5.98/hr | 78.53 | $21.15 |
| NVIDIA B300 | $6.94/hr | 82.9 | $23.25 |
Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.
Speed tiers for Qwen2.5-Coder 32B. 30+ tok/s: 8 (B300, B200, H100 80GB HBM3); 10-30 tok/s: 2 (A10G, L4). 30 tok/s is roughly where replies outpace reading.
Reading your prompt. Before Qwen2.5-Coder 32B writes anything it reads the input: 3417.2 tok/s on the RTX PRO 6000 Blackwell Workstation Edition (1.2s for a 4,000-token prompt), 702.6 on the L4 (5.7s). Long documents and big code files feel this number more than the generation speed.
VRAM for Qwen2.5-Coder 32B. Measured peak 19.2GB, so 24GB is the smallest common card size; smallest card it ran on: A10G (24GB). With long context: Q4_K_M 21GB (tested), Q2_K 14GB, Q3_K_M 17GB, Q5_K_M 25GB, Q6_K 29GB.
Power on Qwen2.5-Coder 32B. Most efficient: H100 80GB HBM3, 175W, 0.64 kWh per 1M generated tokens. Hungriest: B200, 363W, 1.28 kWh. At $0.15/kWh: $0.096 per 1M generated tokens.
Qwen2.5-Coder 32B: 83 tok/s peak, ~21GB floor, a dense coding brain that any 24GB card can hold. Our advice: run it as the strong second opinion in a consensus setup alongside the faster Qwen3 Coder MoE, and don't rent datacenter silicon for single-stream Q4 inference, a 24GB consumer card does half the B300's speed at a sliver of the cost.