QwQ 32B · 10 GPUs measured first-party · llama.cpp Q4_K_M · Updated October 2026
QwQ 32B is Qwen's reasoning model: it thinks out loud before answering, which multiplies token output per question and makes generation speed matter more than for any normal model. We measured it on 10 GPUs (llama.cpp, Q4_K_M): 83 tok/s on the B300, ~21GB peak VRAM, same memory class as the dense coder.
Benchmarked weights: Qwen/QwQ-32B-GGUF

82.91 tok/s on QwQ 32B, the ceiling. Measured on our bench. 288GB of VRAM, $40,000 at launch.

66.07 tok/s on QwQ 32B, lowest launch price that still fits. Measured on our bench. 96GB of VRAM, $8,565 at launch.
What GPU Do You Need for QwQ 32B?, tok/s by GPU
Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.
Efficiency: tok/s per 100W drawn
Power is the average pulled during the run, sampled at 1Hz. The fastest card is often not the one here, and for anything left running this is the number that shows up on the bill.
Value: tok/s per $1,000 of MSRP
Launch price, not street price, so it ages. A speed leaderboard always crowns the most expensive card; this is the counterweight.
QwQ 32B. Measured generation speed by GPU
| GPU | tok/s | Prompt t/s | tok/W | Avg power |
|---|---|---|---|---|
| NVIDIA B300 | 82.91 | 1351.6 | 0.25 | 327.8 W |
| NVIDIA B200 | 78.47 | 2576.7 | 0.22 | 348.9 W |
| NVIDIA H100 80GB HBM3 | 75.82 | 2291.2 | 0.36 | 212.7 W |
| NVIDIA H200 | 75.77 | 2309.1 | 0.34 | 224.6 W |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 66.07 | 3431.8 | 0.3 | 218.3 W |
| NVIDIA A100 80GB SXM4 | 43.27 | 1179.6 | 0.27 | 159.5 W |
| NVIDIA A100 40GB SXM4 | 43.02 | 1155.8 | 0.2 | 220.2 W |
| NVIDIA L40S | 34.59 | 2372.6 | 0.14 | 253.1 W |
| NVIDIA A10G | 21.87 | 797.7 | 0.17 | 132.3 W |
| NVIDIA L4 | 12.28 | 697.3 | 0.18 | 67.4 W |
Reasoning changes the economics. A reasoning model doesn't just answer, it burns hundreds or thousands of 'thinking' tokens first. At 83 tok/s peak, a hard question with a long trace takes real wall-clock time, and on a 23 tok/s card it takes minutes. So the usual 'anything above 30 tok/s feels fine' rule doesn't apply here: with QwQ, throughput is the user experience. Plan for that when choosing hardware. This is the one 32B where the speed column should drive your decision.
Why I still run it. The reasoning does make a difference: on math, logic and tricky debugging, the thinking trace catches errors a direct-answer model confidently commits. My favorite use is in a consensus setup: same prompt to QwQ, a dense coder, and a model from outside the Qwen family, then reconcile. The disagreements are where the insight is. All of them fit a 24GB card one at a time (~21GB floor here), so a single used 3090 or a 7900 XTX can host the whole rotation, no datacenter rental required for single-stream Q4 work.
How it compares. H100 80GB HBM3: QwQ 32B 75.82 tok/s, Qwen2.5-Coder 32B 75.81 (33B), DeepSeek-R1 Distill 32B 75.8 (33B), Olmo-3.1-32B-Think 75.53, Gemma 4 31B 76.59 (31B). 1 of 4 beat QwQ 32B here.
Cost on a rented GPU. 1M generated tokens of QwQ 32B: $3.05 on a A100 40GB SXM4 ($0.47/hr, 6.5 hours), $23.25 on a B300 ($6.94/hr, 3.4 hours, 7.6x the cost).
QwQ 32B: cost per 1M generated tokens on rented GPUs
| GPU | Cheapest rate | Speed (tok/s) | Cost per 1M generated tokens |
|---|---|---|---|
| NVIDIA A100 40GB SXM4 | $0.47/hr | 43.02 | $3.05 |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | $1.08/hr | 66.07 | $4.52 |
| NVIDIA A100 80GB SXM4 | $0.95/hr | 43.27 | $6.08 |
| NVIDIA L40S | $0.79/hr | 34.59 | $6.34 |
| NVIDIA H100 80GB HBM3 | $2.14/hr | 75.82 | $7.83 |
| NVIDIA L4 | $0.44/hr | 12.28 | $9.95 |
| NVIDIA H200 | $3.59/hr | 75.77 | $13.16 |
| NVIDIA B200 | $5.98/hr | 78.47 | $21.17 |
| NVIDIA B300 | $6.94/hr | 82.91 | $23.25 |
Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.
Speed tiers for QwQ 32B. 30+ tok/s: 8 (B300, B200, H100 80GB HBM3); 10-30 tok/s: 2 (A10G, L4). 30 tok/s is roughly where replies outpace reading.
Reading your prompt. Before QwQ 32B writes anything it reads the input: 3431.8 tok/s on the RTX PRO 6000 Blackwell Workstation Edition (1.2s for a 4,000-token prompt), 697.3 on the L4 (5.7s). Long documents and big code files feel this number more than the generation speed.
VRAM for QwQ 32B. Measured peak 19.2GB, so 24GB is the smallest common card size; smallest card it ran on: A10G (24GB). With long context: Q4_K_M 21GB (tested), Q2_K 14GB, Q3_K_M 17GB, Q5_K_M 25GB, Q6_K 29GB.
Power on QwQ 32B. Most efficient: H100 80GB HBM3, 213W, 0.78 kWh per 1M generated tokens. Hungriest: B200, 349W, 1.24 kWh. At $0.15/kWh: $0.12 per 1M generated tokens.
QwQ 32B: 83 tok/s peak, ~21GB floor, and a token-hungry reasoning style that makes speed the whole ballgame. Worth it when the problem is genuinely hard, and at its best as one voice in a multi-model consensus rather than your only local brain.