QwQ 32B · 10 GPUs measured first-party · llama.cpp Q4_K_M · Updated October 2026

What GPU Do You Need for QwQ 32B?

QwQ 32B is Qwen's reasoning model: it thinks out loud before answering, which multiplies token output per question and makes generation speed matter more than for any normal model. We measured it on 10 GPUs (llama.cpp, Q4_K_M): 83 tok/s on the B300, ~21GB peak VRAM, same memory class as the dense coder.

Benchmarked weights: Qwen/QwQ-32B-GGUF

Fastest we measured
NVIDIA B300

NVIDIA B300

82.91 tok/s on QwQ 32B, the ceiling. Measured on our bench. 288GB of VRAM, $40,000 at launch.

Pros
  • 82.91 tok/s on QwQ 32B
  • 288GB, clears the QwQ 32B floor
  • Rentable by the hour rather than bought
Cons
  • 1400W board rating
  • Datacenter or workstation hardware, not a retail purchase
Cheapest card that runs it
NVIDIA RTX PRO 6000 Blackwell Workstation Edition

NVIDIA RTX PRO 6000 Blackwell Workstation Edition

66.07 tok/s on QwQ 32B, lowest launch price that still fits. Measured on our bench. 96GB of VRAM, $8,565 at launch.

Pros
  • 66.07 tok/s on QwQ 32B
  • 96GB, clears the QwQ 32B floor
  • Rentable by the hour rather than bought
Cons
  • 600W board rating
  • Datacenter or workstation hardware, not a retail purchase
82.91tok/s
Fastest: NVIDIA B300
measured
10
Cards that run QwQ 32B
of 11 we have data for
1
Cards that can't run it at all
published as hard gates, not omissions
575%
Fastest vs slowest that fits
82.91 vs 12.28 tok/s

What GPU Do You Need for QwQ 32B?, tok/s by GPU

NVIDIA B300
82.91 tok/s
NVIDIA B200
78.47 tok/s
NVIDIA H100 80GB HBM3
75.82 tok/s
NVIDIA H200
75.77 tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
66.07 tok/s
NVIDIA A100 80GB SXM4
43.27 tok/s
NVIDIA A100 40GB SXM4
43.02 tok/s
NVIDIA L40S
34.59 tok/s
NVIDIA A10G
21.87 tok/s
NVIDIA L4
12.28 tok/s

Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.

Efficiency: tok/s per 100W drawn

NVIDIA H100 80GB HBM3
35.65 tok/s / 100W
NVIDIA H200
33.74 tok/s / 100W
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
30.27 tok/s / 100W
NVIDIA A100 80GB SXM4
27.13 tok/s / 100W
NVIDIA B300
25.29 tok/s / 100W
NVIDIA B200
22.49 tok/s / 100W
NVIDIA A100 40GB SXM4
19.54 tok/s / 100W
NVIDIA L4
18.22 tok/s / 100W
NVIDIA A10G
16.53 tok/s / 100W
NVIDIA L40S
13.67 tok/s / 100W

Power is the average pulled during the run, sampled at 1Hz. The fastest card is often not the one here, and for anything left running this is the number that shows up on the bill.

Value: tok/s per $1,000 of MSRP

NVIDIA A10G
7.81 tok/s / $1k
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
7.71 tok/s / $1k
NVIDIA L4
4.91 tok/s / $1k
NVIDIA L40S
4.61 tok/s / $1k
NVIDIA A100 40GB SXM4
3.58 tok/s / $1k
NVIDIA A100 80GB SXM4
2.55 tok/s / $1k
NVIDIA H100 80GB HBM3
2.53 tok/s / $1k
NVIDIA H200
2.44 tok/s / $1k
NVIDIA B300
2.07 tok/s / $1k
NVIDIA B200
1.96 tok/s / $1k

Launch price, not street price, so it ages. A speed leaderboard always crowns the most expensive card; this is the counterweight.

QwQ 32B. Measured generation speed by GPU

NVIDIA B30082.91
NVIDIA B20078.47
NVIDIA H100 80GB HBM375.82
NVIDIA H20075.77
NVIDIA RTX PRO 6000 Blackwell Workstation Edition66.07
NVIDIA A100 80GB SXM443.27
NVIDIA A100 40GB SXM443.02
NVIDIA L40S34.59
NVIDIA A10G21.87
NVIDIA L412.28
GPUtok/sPrompt t/stok/WAvg power
NVIDIA B30082.911351.60.25327.8 W
NVIDIA B20078.472576.70.22348.9 W
NVIDIA H100 80GB HBM375.822291.20.36212.7 W
NVIDIA H20075.772309.10.34224.6 W
NVIDIA RTX PRO 6000 Blackwell Workstation Edition66.073431.80.3218.3 W
NVIDIA A100 80GB SXM443.271179.60.27159.5 W
NVIDIA A100 40GB SXM443.021155.80.2220.2 W
NVIDIA L40S34.592372.60.14253.1 W
NVIDIA A10G21.87797.70.17132.3 W
NVIDIA L412.28697.30.1867.4 W

Reasoning changes the economics. A reasoning model doesn't just answer, it burns hundreds or thousands of 'thinking' tokens first. At 83 tok/s peak, a hard question with a long trace takes real wall-clock time, and on a 23 tok/s card it takes minutes. So the usual 'anything above 30 tok/s feels fine' rule doesn't apply here: with QwQ, throughput is the user experience. Plan for that when choosing hardware. This is the one 32B where the speed column should drive your decision.

Why I still run it. The reasoning does make a difference: on math, logic and tricky debugging, the thinking trace catches errors a direct-answer model confidently commits. My favorite use is in a consensus setup: same prompt to QwQ, a dense coder, and a model from outside the Qwen family, then reconcile. The disagreements are where the insight is. All of them fit a 24GB card one at a time (~21GB floor here), so a single used 3090 or a 7900 XTX can host the whole rotation, no datacenter rental required for single-stream Q4 work.

How it compares. H100 80GB HBM3: QwQ 32B 75.82 tok/s, Qwen2.5-Coder 32B 75.81 (33B), DeepSeek-R1 Distill 32B 75.8 (33B), Olmo-3.1-32B-Think 75.53, Gemma 4 31B 76.59 (31B). 1 of 4 beat QwQ 32B here.

Cost on a rented GPU. 1M generated tokens of QwQ 32B: $3.05 on a A100 40GB SXM4 ($0.47/hr, 6.5 hours), $23.25 on a B300 ($6.94/hr, 3.4 hours, 7.6x the cost).

QwQ 32B: cost per 1M generated tokens on rented GPUs

NVIDIA A100 40GB SXM4$0.47/hr
NVIDIA RTX PRO 6000 Blackwell Workstation Edition$1.08/hr
NVIDIA A100 80GB SXM4$0.95/hr
NVIDIA L40S$0.79/hr
NVIDIA H100 80GB HBM3$2.14/hr
NVIDIA L4$0.44/hr
NVIDIA H200$3.59/hr
NVIDIA B200$5.98/hr
NVIDIA B300$6.94/hr
GPUCheapest rateSpeed (tok/s)Cost per 1M generated tokens
NVIDIA A100 40GB SXM4$0.47/hr43.02$3.05
NVIDIA RTX PRO 6000 Blackwell Workstation Edition$1.08/hr66.07$4.52
NVIDIA A100 80GB SXM4$0.95/hr43.27$6.08
NVIDIA L40S$0.79/hr34.59$6.34
NVIDIA H100 80GB HBM3$2.14/hr75.82$7.83
NVIDIA L4$0.44/hr12.28$9.95
NVIDIA H200$3.59/hr75.77$13.16
NVIDIA B200$5.98/hr78.47$21.17
NVIDIA B300$6.94/hr82.91$23.25

Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.

Speed tiers for QwQ 32B. 30+ tok/s: 8 (B300, B200, H100 80GB HBM3); 10-30 tok/s: 2 (A10G, L4). 30 tok/s is roughly where replies outpace reading.

Reading your prompt. Before QwQ 32B writes anything it reads the input: 3431.8 tok/s on the RTX PRO 6000 Blackwell Workstation Edition (1.2s for a 4,000-token prompt), 697.3 on the L4 (5.7s). Long documents and big code files feel this number more than the generation speed.

VRAM for QwQ 32B. Measured peak 19.2GB, so 24GB is the smallest common card size; smallest card it ran on: A10G (24GB). With long context: Q4_K_M 21GB (tested), Q2_K 14GB, Q3_K_M 17GB, Q5_K_M 25GB, Q6_K 29GB.

Power on QwQ 32B. Most efficient: H100 80GB HBM3, 213W, 0.78 kWh per 1M generated tokens. Hungriest: B200, 349W, 1.24 kWh. At $0.15/kWh: $0.12 per 1M generated tokens.

Our verdict

QwQ 32B: 83 tok/s peak, ~21GB floor, and a token-hungry reasoning style that makes speed the whole ballgame. Worth it when the problem is genuinely hard, and at its best as one voice in a multi-model consensus rather than your only local brain.

FAQ

Why does speed matter more for QwQ than other models?
It generates long thinking traces before every answer, often thousands of tokens. At 83 tok/s that's tolerable; on a slow card it's minutes per question. Judge cards for QwQ by the tok/s column above everything else.
What GPU fits QwQ 32B?
24GB cards and up. Measured peak was ~21GB at Q4_K_M. RX 7900 XTX, RTX 3090, RTX 4090, or rented datacenter cards (B300 topped our chart at 83 tok/s).
Is the reasoning actually worth the token cost?
On hard problems, math, logic, subtle bugs, genuinely yes: the trace catches mistakes direct-answer models commit to. On easy questions it's pure overhead. Route accordingly: QwQ for the hard 10%, a fast model for the rest.
QwQ 32B or Qwen3 32B?
Identical hardware requirements and near-identical measured speed (83 vs 84 tok/s at the top). QwQ trades latency for deliberation; Qwen3 32B answers directly. Many setups keep both and route by difficulty.
What's the best way to use QwQ locally?
In a consensus rotation: same prompt to QwQ, a coder model, and a non-Qwen 32B, then compare. One 24GB card hosts all three (one at a time), and the disagreement between them is where the value shows up.