Qwen2.5-Coder 32B · 10 GPUs measured first-party · llama.cpp Q4_K_M · Updated July 2026
Qwen2.5-Coder 32B was the local coding flagship of its generation: a dense 32B tuned for code, needing ~21GB at Q4_K_M and topping our chart at 83 tok/s on the B300. We measured it on 10 GPUs with the same pinned llama.cpp harness as every other model on this site.
Benchmarked weights: bartowski/Qwen2.5-Coder-32B-Instruct-GGUF
What GPU Do You Need for Qwen2.5-Coder 32B?, tok/s, fastest 10
Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.
Qwen2.5-Coder 32B. Measured generation speed by GPU
| GPU | tok/s | Prompt t/s | tok/W | Avg power |
|---|---|---|---|---|
| NVIDIA B300 | 82.9 | 1336.7 | 0.26 | 316.1 W |
| NVIDIA B200 | 78.53 | 2558.4 | 0.22 | 362.8 W |
| NVIDIA H100 80GB HBM3 | 75.81 | 2295.6 | 0.43 | 174.7 W |
| NVIDIA H200 | 75.77 | 2289.7 | 0.29 | 259.8 W |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 66.07 | 3417.2 | 0.28 | 239.3 W |
| NVIDIA A100 40GB SXM4 | 43.54 | 1174.7 | 0.28 | 155.6 W |
| NVIDIA A100 80GB SXM4 | 43.36 | 1188.4 | 0.25 | 174.0 W |
| NVIDIA L40S | 34.46 | 2286.4 | 0.18 | 189.2 W |
| NVIDIA A10G | 23.52 | 1036.1 | 0.16 | 150.6 W |
| NVIDIA L4 | 11.82 | 666.3 | 0.2 | 60.5 W |
Where a dense 32B coder fits now. Two ways to use this model well. First, as the strong single model on a 24GB card: 24GB clears the ~21GB floor, and 30-40 tok/s-class generation is comfortable for chat-style coding. Second, and more interesting: as one voice in a consensus setup. Running the same prompt through two or three different 32B-class models, this, QwQ, something outside the Qwen family, and reconciling the answers gets you meaningfully better results than any single model. That's my current view of the best local strategy: dense 32B models as ensemble members, not soloists. Since the Qwen3 Coder MoE arrived, the MoE is the daily driver and this is the second opinion.
A value note the chart hides. The B300 tops the board at 83 tok/s, but dense 32B generation is bandwidth-bound and single-stream, so datacenter silicon is mostly wasted on it. On the sister dense model (Qwen3 32B) we measured an RTX 5090 at 71 tok/s against the B300's 84: a consumer card delivering 85% of the flagship datacenter chip's output on this workload. Even a used RTX 3090 holds 38 tok/s. Renting big iron for one-at-a-time Q4 inference is the least efficient way to spend AI money, big cards earn their rate on batch serving, not solo chat.
Qwen2.5-Coder 32B: 83 tok/s peak, ~21GB floor, a dense coding brain that any 24GB card can hold. Our advice: run it as the strong second opinion in a consensus setup alongside the faster Qwen3 Coder MoE, and don't rent datacenter silicon for single-stream Q4 inference, a 24GB consumer card does half the B300's speed at a sliver of the cost.