DeepSeek-R1 Distill 7B · 11 GPUs measured first-party · llama.cpp Q4_K_M · Updated October 2026

What GPU Do You Need for DeepSeek-R1 Distill 7B?

DeepSeek-R1 Distill 7B is where this family starts being worth your VRAM: in our view, the minimum R1 distill for real work. Measured on 11 GPUs (llama.cpp, Q4_K_M): 285 tok/s on the B300, ~6GB peak, which puts genuine reasoning within reach of any 8GB card.

Benchmarked weights: bartowski/DeepSeek-R1-Distill-Qwen-7B-GGUF

Fastest we measured
NVIDIA B300

NVIDIA B300

284.8 tok/s on DeepSeek-R1 Distill 7B, the ceiling. Measured on our bench. 288GB of VRAM, $40,000 at launch.

Pros
  • 284.8 tok/s on DeepSeek-R1 Distill 7B
  • 288GB, clears the DeepSeek-R1 Distill 7B floor
  • Rentable by the hour rather than bought
Cons
  • 1400W board rating
  • Datacenter or workstation hardware, not a retail purchase
Cheapest card that runs it
NVIDIA RTX PRO 6000 Blackwell Workstation Edition

NVIDIA RTX PRO 6000 Blackwell Workstation Edition

256.0 tok/s on DeepSeek-R1 Distill 7B, lowest launch price that still fits. Measured on our bench. 96GB of VRAM, $8,565 at launch.

Pros
  • 256.0 tok/s on DeepSeek-R1 Distill 7B
  • 96GB, clears the DeepSeek-R1 Distill 7B floor
  • Rentable by the hour rather than bought
Cons
  • 600W board rating
  • Datacenter or workstation hardware, not a retail purchase
284.8tok/s
Fastest: NVIDIA B300
measured
11
Cards that run DeepSeek-R1 Distill 7B
of 11 we have data for
0
Cards that can't run it at all
published as hard gates, not omissions
658%
Fastest vs slowest that fits
284.8 vs 37.58 tok/s

What GPU Do You Need for DeepSeek-R1 Distill 7B?, tok/s by GPU

NVIDIA B300
284.8 tok/s
NVIDIA B200
277 tok/s
NVIDIA H200
267.1 tok/s
NVIDIA H100 80GB HBM3
264.4 tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
256 tok/s
NVIDIA A100 80GB SXM4
163 tok/s
NVIDIA A100 40GB SXM4
159.1 tok/s
NVIDIA L40S
144 tok/s
NVIDIA A10G
89.79 tok/s
NVIDIA L4
53.27 tok/s
NVIDIA T4
37.58 tok/s

Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.

Efficiency: tok/s per 100W drawn

NVIDIA H200
243.48 tok/s / 100W
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
151.5 tok/s / 100W
NVIDIA H100 80GB HBM3
123.97 tok/s / 100W
NVIDIA A100 80GB SXM4
106.58 tok/s / 100W
NVIDIA B300
97.99 tok/s / 100W
NVIDIA A100 40GB SXM4
96.83 tok/s / 100W
NVIDIA B200
92.71 tok/s / 100W
NVIDIA L4
84.96 tok/s / 100W
NVIDIA A10G
73.72 tok/s / 100W
NVIDIA L40S
72.38 tok/s / 100W
NVIDIA T4
59.46 tok/s / 100W

Power is the average pulled during the run, sampled at 1Hz. The fastest card is often not the one here, and for anything left running this is the number that shows up on the bill.

Value: tok/s per $1,000 of MSRP

NVIDIA A10G
32.07 tok/s / $1k
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
29.89 tok/s / $1k
NVIDIA L4
21.31 tok/s / $1k
NVIDIA L40S
19.2 tok/s / $1k
NVIDIA T4
16.35 tok/s / $1k
NVIDIA A100 40GB SXM4
13.26 tok/s / $1k
NVIDIA A100 80GB SXM4
9.59 tok/s / $1k
NVIDIA H100 80GB HBM3
8.81 tok/s / $1k
NVIDIA H200
8.62 tok/s / $1k
NVIDIA B300
7.12 tok/s / $1k
NVIDIA B200
6.93 tok/s / $1k

Launch price, not street price, so it ages. A speed leaderboard always crowns the most expensive card; this is the counterweight.

DeepSeek-R1 Distill 7B. Measured generation speed by GPU

NVIDIA B300284.8
NVIDIA B200277
NVIDIA H200267.1
NVIDIA H100 80GB HBM3264.4
NVIDIA RTX PRO 6000 Blackwell Workstation Edition256
NVIDIA A100 80GB SXM4163
NVIDIA A100 40GB SXM4159.1
NVIDIA L40S144
NVIDIA A10G89.79
NVIDIA L453.27
NVIDIA T437.58
GPUtok/sPrompt t/stok/WAvg power
NVIDIA B300284.85662.40.98290.6 W
NVIDIA B20027710323.10.93298.8 W
NVIDIA H200267.19028.62.43109.7 W
NVIDIA H100 80GB HBM3264.49347.71.24213.3 W
NVIDIA RTX PRO 6000 Blackwell Workstation Edition25612909.71.52169.0 W
NVIDIA A100 80GB SXM41634844.41.07152.9 W
NVIDIA A100 40GB SXM4159.14649.90.97164.3 W
NVIDIA L40S14410096.30.72198.9 W
NVIDIA A10G89.793506.30.74121.8 W
NVIDIA L453.273197.90.8562.7 W
NVIDIA T437.581335.60.5963.2 W

The entry ticket to local reasoning. Below 7B, R1-style thinking traces are mostly noise; from 7B up they start catching real mistakes. That makes this the cheapest model we'd trust with reasoning-flavored work, and at a ~6GB floor, the hardware bar is an 8GB card, not a workstation. It's also a natural junior member in a consensus setup: run it alongside a Qwen model of similar size and compare answers; where they disagree is usually where the problem is interesting.

Reading the chart. The top is tight, B300 at 285, B200 at 277, H200 at 267 tok/s, but the efficiency column isn't: the H200 does its 267 at 110W (2.43 tok/W), roughly two and a half times the efficiency of either Blackwell card. And because reasoning models burn tokens thinking, sustained throughput per watt matters more here than for direct-answer models of the same size. At the budget end, 38 tok/s on a T4 is still usable for a single patient user.

How it compares. H100 80GB HBM3: DeepSeek-R1 Distill 7B 264.4 tok/s, Llama 3 8B 264.4 (8B), Qwen2.5-Coder-7B-Instruct-abliterated 264.2, Qwen2.5-7B 264.2 (8B), Qwen2.5-Coder 7B 263.9 (8B). DeepSeek-R1 Distill 7B beats all 4 here.

Cost on a rented GPU. 1M generated tokens of DeepSeek-R1 Distill 7B: $0.82 on a A100 40GB SXM4 ($0.47/hr, 105 min), $6.77 on a B300 ($6.94/hr, 59 min, 8.2x the cost).

DeepSeek-R1 Distill 7B: cost per 1M generated tokens on rented GPUs

NVIDIA A100 40GB SXM4$0.47/hr
NVIDIA T4$0.14/hr
NVIDIA RTX PRO 6000 Blackwell Workstation Edition$1.08/hr
NVIDIA L40S$0.79/hr
NVIDIA A100 80GB SXM4$0.95/hr
NVIDIA H100 80GB HBM3$2.14/hr
NVIDIA L4$0.44/hr
NVIDIA H200$3.59/hr
NVIDIA B200$5.98/hr
NVIDIA B300$6.94/hr
GPUCheapest rateSpeed (tok/s)Cost per 1M generated tokens
NVIDIA A100 40GB SXM4$0.47/hr159.1$0.82
NVIDIA T4$0.14/hr37.58$1.01
NVIDIA RTX PRO 6000 Blackwell Workstation Edition$1.08/hr256$1.17
NVIDIA L40S$0.79/hr144$1.52
NVIDIA A100 80GB SXM4$0.95/hr163$1.61
NVIDIA H100 80GB HBM3$2.14/hr264.4$2.24
NVIDIA L4$0.44/hr53.27$2.29
NVIDIA H200$3.59/hr267.1$3.73
NVIDIA B200$5.98/hr277$6.00
NVIDIA B300$6.94/hr284.8$6.77

Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.

Speed tiers for DeepSeek-R1 Distill 7B. 30+ tok/s: 11 (B300, B200, H200). 30 tok/s is roughly where replies outpace reading.

Reading your prompt. Before DeepSeek-R1 Distill 7B writes anything it reads the input: 12909.7 tok/s on the RTX PRO 6000 Blackwell Workstation Edition (0.3s for a 4,000-token prompt), 1335.6 on the T4 (3.0s). Long documents and big code files feel this number more than the generation speed.

VRAM for DeepSeek-R1 Distill 7B. Measured peak 5.0GB, so 8GB is the smallest common card size; smallest card it ran on: T4 (16GB). With long context: Q4_K_M 6GB (tested), Q2_K 4GB, Q3_K_M 5GB, Q5_K_M 7GB, Q6_K 9GB.

Power on DeepSeek-R1 Distill 7B. Most efficient: RTX PRO 6000 Blackwell Workstation Edition, 169W, 0.18 kWh per 1M generated tokens. Hungriest: B200, 299W, 0.30 kWh. At $0.15/kWh: $0.028 per 1M generated tokens.

Our verdict

DeepSeek-R1 Distill 7B: 285 tok/s peak, ~6GB floor, the smallest R1 distill whose reasoning actually earns its token burn. Our pick for the cheapest genuine reasoning setup: this on any 8GB card, with a same-size Qwen as its second opinion.

FAQ

What GPU do I need for DeepSeek-R1 Distill 7B?
Any 8GB card. Measured peak was ~6GB at Q4_K_M. The $179 Intel Arc A580 clears it, and every datacenter card we tested runs it comfortably.
Is 7B the right size for a reasoning model?
It's the minimum where the thinking traces reliably help rather than just costing tokens. Below it (the 1.5B distill), reasoning is mostly ceremony; above it (14B/32B), quality scales with your VRAM.
How does it pair with Qwen models?
Very well. That's our recommended use. DeepSeek and Qwen have different strengths by field, so running both on the same prompt and comparing answers catches errors either one makes alone. Both 7B-class models fit an 8GB card one at a time.
Why do reasoning models need more tok/s than normal models?
They generate thinking tokens before every answer, sometimes thousands. At 267 tok/s on an H200 that's seconds; at 38 tok/s on a T4 it's a coffee break. Judge hardware for this family by the speed column first.
Distill-Qwen 7B or Distill-Llama 8B?
Nearly identical speed and VRAM in our tests (285 vs 284 tok/s peak, both ~6GB). The base model differs, Qwen vs Llama, so behavior differs by task. For ensembles, using one of each is exactly the kind of diversity that makes consensus work.