DeepSeek-R1 Distill 32B · 10 GPUs measured first-party · llama.cpp Q4_K_M · Updated July 2026
DeepSeek-R1 Distill 32B is the flagship of the distill family, the closest you get to full R1-style reasoning in a package a 24GB card can hold. Measured on 10 GPUs (llama.cpp, Q4_K_M): 83 tok/s on the B300, ~21GB peak VRAM.
Benchmarked weights: bartowski/DeepSeek-R1-Distill-Qwen-32B-GGUF
What GPU Do You Need for DeepSeek-R1 Distill 32B?, tok/s, fastest 10
Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.
DeepSeek-R1 Distill 32B. Measured generation speed by GPU
| GPU | tok/s | Prompt t/s | tok/W | Avg power |
|---|---|---|---|---|
| NVIDIA B300 | 82.9 | 1363 | 0.24 | 341.7 W |
| NVIDIA B200 | 78.55 | 2593.9 | 0.2 | 392.3 W |
| NVIDIA H100 80GB HBM3 | 75.8 | 2280.5 | 0.36 | 209.8 W |
| NVIDIA H200 | 75.8 | 2257.9 | 0.31 | 243.9 W |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 66.06 | 3428.6 | 0.25 | 260.1 W |
| NVIDIA A100 40GB SXM4 | 43.55 | 1166.1 | 0.27 | 160.5 W |
| NVIDIA A100 80GB SXM4 | 43.11 | 1181.3 | 0.23 | 185.6 W |
| NVIDIA L40S | 34.45 | 2332.1 | 0.18 | 194.7 W |
| NVIDIA A10G | 23.5 | 1032.5 | 0.16 | 146.9 W |
| NVIDIA L4 | 11.78 | 664.5 | 0.19 | 60.9 W |
The pairing we actually recommend. Here's the play this model was born for: DeepSeek 32B plus a Qwen 32B-class model on the same prompt, two strong brains from different labs, each double-checking the other. That two-model setup produces higher-quality output than any single model either family offers, and one 24GB card hosts the whole rotation. It's the strongest version of our consensus argument: don't hunt for the one biggest model you can fit; assemble models that are better in different fields and make them audit each other.
Hardware value, measured. Dense 32B inference is bandwidth-bound and single-stream, and the chart shows what that means for spending: the B300 tops out at 83 tok/s while, on the sister dense Qwen3 32B, an RTX 5090 measured 71 tok/s, 85% of the datacenter flagship from a consumer card. Rent big iron for batch serving; for a personal reasoning panel, a 5090, or a 24GB card at half the speed, is the rational buy. Budget for the trace tax too: at 23 tok/s (A10G) a long think is a minute of waiting.
DeepSeek-R1 Distill 32B: 83 tok/s peak, ~21GB floor, and the best reason we know to own a big consumer card, pair it with a Qwen 32B in a two-model consensus and you get output quality neither achieves alone, all on hardware that sits under your desk.