DeepSeek-R1 Distill 7B · 11 GPUs measured first-party · llama.cpp Q4_K_M · Updated October 2026
DeepSeek-R1 Distill 7B is where this family starts being worth your VRAM: in our view, the minimum R1 distill for real work. Measured on 11 GPUs (llama.cpp, Q4_K_M): 285 tok/s on the B300, ~6GB peak, which puts genuine reasoning within reach of any 8GB card.
Benchmarked weights: bartowski/DeepSeek-R1-Distill-Qwen-7B-GGUF

284.8 tok/s on DeepSeek-R1 Distill 7B, the ceiling. Measured on our bench. 288GB of VRAM, $40,000 at launch.

256.0 tok/s on DeepSeek-R1 Distill 7B, lowest launch price that still fits. Measured on our bench. 96GB of VRAM, $8,565 at launch.
What GPU Do You Need for DeepSeek-R1 Distill 7B?, tok/s by GPU
Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.
Efficiency: tok/s per 100W drawn
Power is the average pulled during the run, sampled at 1Hz. The fastest card is often not the one here, and for anything left running this is the number that shows up on the bill.
Value: tok/s per $1,000 of MSRP
Launch price, not street price, so it ages. A speed leaderboard always crowns the most expensive card; this is the counterweight.
DeepSeek-R1 Distill 7B. Measured generation speed by GPU
| GPU | tok/s | Prompt t/s | tok/W | Avg power |
|---|---|---|---|---|
| NVIDIA B300 | 284.8 | 5662.4 | 0.98 | 290.6 W |
| NVIDIA B200 | 277 | 10323.1 | 0.93 | 298.8 W |
| NVIDIA H200 | 267.1 | 9028.6 | 2.43 | 109.7 W |
| NVIDIA H100 80GB HBM3 | 264.4 | 9347.7 | 1.24 | 213.3 W |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 256 | 12909.7 | 1.52 | 169.0 W |
| NVIDIA A100 80GB SXM4 | 163 | 4844.4 | 1.07 | 152.9 W |
| NVIDIA A100 40GB SXM4 | 159.1 | 4649.9 | 0.97 | 164.3 W |
| NVIDIA L40S | 144 | 10096.3 | 0.72 | 198.9 W |
| NVIDIA A10G | 89.79 | 3506.3 | 0.74 | 121.8 W |
| NVIDIA L4 | 53.27 | 3197.9 | 0.85 | 62.7 W |
| NVIDIA T4 | 37.58 | 1335.6 | 0.59 | 63.2 W |
The entry ticket to local reasoning. Below 7B, R1-style thinking traces are mostly noise; from 7B up they start catching real mistakes. That makes this the cheapest model we'd trust with reasoning-flavored work, and at a ~6GB floor, the hardware bar is an 8GB card, not a workstation. It's also a natural junior member in a consensus setup: run it alongside a Qwen model of similar size and compare answers; where they disagree is usually where the problem is interesting.
Reading the chart. The top is tight, B300 at 285, B200 at 277, H200 at 267 tok/s, but the efficiency column isn't: the H200 does its 267 at 110W (2.43 tok/W), roughly two and a half times the efficiency of either Blackwell card. And because reasoning models burn tokens thinking, sustained throughput per watt matters more here than for direct-answer models of the same size. At the budget end, 38 tok/s on a T4 is still usable for a single patient user.
How it compares. H100 80GB HBM3: DeepSeek-R1 Distill 7B 264.4 tok/s, Llama 3 8B 264.4 (8B), Qwen2.5-Coder-7B-Instruct-abliterated 264.2, Qwen2.5-7B 264.2 (8B), Qwen2.5-Coder 7B 263.9 (8B). DeepSeek-R1 Distill 7B beats all 4 here.
Cost on a rented GPU. 1M generated tokens of DeepSeek-R1 Distill 7B: $0.82 on a A100 40GB SXM4 ($0.47/hr, 105 min), $6.77 on a B300 ($6.94/hr, 59 min, 8.2x the cost).
DeepSeek-R1 Distill 7B: cost per 1M generated tokens on rented GPUs
| GPU | Cheapest rate | Speed (tok/s) | Cost per 1M generated tokens |
|---|---|---|---|
| NVIDIA A100 40GB SXM4 | $0.47/hr | 159.1 | $0.82 |
| NVIDIA T4 | $0.14/hr | 37.58 | $1.01 |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | $1.08/hr | 256 | $1.17 |
| NVIDIA L40S | $0.79/hr | 144 | $1.52 |
| NVIDIA A100 80GB SXM4 | $0.95/hr | 163 | $1.61 |
| NVIDIA H100 80GB HBM3 | $2.14/hr | 264.4 | $2.24 |
| NVIDIA L4 | $0.44/hr | 53.27 | $2.29 |
| NVIDIA H200 | $3.59/hr | 267.1 | $3.73 |
| NVIDIA B200 | $5.98/hr | 277 | $6.00 |
| NVIDIA B300 | $6.94/hr | 284.8 | $6.77 |
Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.
Speed tiers for DeepSeek-R1 Distill 7B. 30+ tok/s: 11 (B300, B200, H200). 30 tok/s is roughly where replies outpace reading.
Reading your prompt. Before DeepSeek-R1 Distill 7B writes anything it reads the input: 12909.7 tok/s on the RTX PRO 6000 Blackwell Workstation Edition (0.3s for a 4,000-token prompt), 1335.6 on the T4 (3.0s). Long documents and big code files feel this number more than the generation speed.
VRAM for DeepSeek-R1 Distill 7B. Measured peak 5.0GB, so 8GB is the smallest common card size; smallest card it ran on: T4 (16GB). With long context: Q4_K_M 6GB (tested), Q2_K 4GB, Q3_K_M 5GB, Q5_K_M 7GB, Q6_K 9GB.
Power on DeepSeek-R1 Distill 7B. Most efficient: RTX PRO 6000 Blackwell Workstation Edition, 169W, 0.18 kWh per 1M generated tokens. Hungriest: B200, 299W, 0.30 kWh. At $0.15/kWh: $0.028 per 1M generated tokens.
DeepSeek-R1 Distill 7B: 285 tok/s peak, ~6GB floor, the smallest R1 distill whose reasoning actually earns its token burn. Our pick for the cheapest genuine reasoning setup: this on any 8GB card, with a same-size Qwen as its second opinion.