DeepSeek-R1 Distill 7B · 11 GPUs measured first-party · llama.cpp Q4_K_M · Updated July 2026
DeepSeek-R1 Distill 7B is where this family starts being worth your VRAM: in our view, the minimum R1 distill for real work. Measured on 11 GPUs (llama.cpp, Q4_K_M): 285 tok/s on the B300, ~6GB peak, which puts genuine reasoning within reach of any 8GB card.
Benchmarked weights: bartowski/DeepSeek-R1-Distill-Qwen-7B-GGUF
What GPU Do You Need for DeepSeek-R1 Distill 7B?, tok/s, fastest 11
Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.
DeepSeek-R1 Distill 7B. Measured generation speed by GPU
| GPU | tok/s | Prompt t/s | tok/W | Avg power |
|---|---|---|---|---|
| NVIDIA B300 | 284.77 | 5662.4 | 0.98 | 290.6 W |
| NVIDIA B200 | 277.02 | 10323.1 | 0.93 | 298.8 W |
| NVIDIA H200 | 267.1 | 9028.6 | 2.43 | 109.7 W |
| NVIDIA H100 80GB HBM3 | 264.42 | 9347.7 | 1.24 | 213.3 W |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 256.04 | 12909.7 | 1.52 | 169.0 W |
| NVIDIA A100 80GB SXM4 | 162.96 | 4844.4 | 1.07 | 152.9 W |
| NVIDIA A100 40GB SXM4 | 161.37 | 4751.3 | 1.15 | 140.3 W |
| NVIDIA L40S | 143.13 | 9832.2 | 0.95 | 150.9 W |
| NVIDIA A10G | 98.05 | 4094 | 0.81 | 121.3 W |
| NVIDIA L4 | 53.27 | 3193.6 | 0.97 | 54.9 W |
| NVIDIA T4 | 38.3 | 1349.6 | 0.68 | 56.1 W |
The entry ticket to local reasoning. Below 7B, R1-style thinking traces are mostly noise; from 7B up they start catching real mistakes. That makes this the cheapest model we'd trust with reasoning-flavored work, and at a ~6GB floor, the hardware bar is an 8GB card, not a workstation. It's also a natural junior member in a consensus setup: run it alongside a Qwen model of similar size and compare answers; where they disagree is usually where the problem is interesting.
Reading the chart. The top is tight, B300 at 285, B200 at 277, H200 at 267 tok/s, but the efficiency column isn't: the H200 does its 267 at 110W (2.43 tok/W), roughly two and a half times the efficiency of either Blackwell card. And because reasoning models burn tokens thinking, sustained throughput per watt matters more here than for direct-answer models of the same size. At the budget end, 38 tok/s on a T4 is still usable for a single patient user.
DeepSeek-R1 Distill 7B: 285 tok/s peak, ~6GB floor, the smallest R1 distill whose reasoning actually earns its token burn. Our pick for the cheapest genuine reasoning setup: this on any 8GB card, with a same-size Qwen as its second opinion.