DeepSeek-R1 Distill 14B (Q3_K_M) · 11 GPUs measured first-party · llama.cpp · Updated October 2026

What GPU Do You Need for DeepSeek-R1 Distill 14B (Q3_K_M)?

DeepSeek-R1 Distill 14B (Q3_K_M) on 11 GPUs, measured first-party: NVIDIA RTX PRO 6000 Blackwell Workstation Edition leads at 149.8 tok/s, T4 trails at 16.9 tok/s, and it peaked at 8GB of VRAM.

Benchmarked weights: bartowski/DeepSeek-R1-Distill-Qwen-14B-GGUF

Fastest we measured
NVIDIA RTX PRO 6000 Blackwell Workstation Edition

NVIDIA RTX PRO 6000 Blackwell Workstation Edition

149.8 tok/s on DeepSeek-R1 Distill 14B (Q3_K_M), the ceiling. Measured on our bench. 96GB of VRAM, $8,565 at launch.

Pros
  • 149.8 tok/s on DeepSeek-R1 Distill 14B (Q3_K_M)
  • 96GB, clears the DeepSeek-R1 Distill 14B (Q3_K_M) floor
  • Rentable by the hour rather than bought
Cons
  • 600W board rating
  • Datacenter or workstation hardware, not a retail purchase
149.8tok/s
Fastest: NVIDIA RTX PRO 6000 Blackwell Workstation Edition
measured, 3-run average
~8GB
VRAM needed (measured peak)
GPU-independent, applies to every card
11
GPUs measured
same pinned harness
0.55tok/s/W
Most efficient: NVIDIA RTX PRO 6000 Blackwell Workstation Edition
real power sampling, not TDP

What GPU Do You Need for DeepSeek-R1 Distill 14B (Q3_K_M)?, tok/s by GPU

NVIDIA RTX PRO 6000 Blackwell Workstation Edition
149.8 tok/s
NVIDIA B300
144.7 tok/s
NVIDIA B200
120.5 tok/s
NVIDIA H200
119.7 tok/s
NVIDIA H100 80GB HBM3
118.4 tok/s
NVIDIA L40S
85.85 tok/s
NVIDIA A100 80GB SXM4
67.13 tok/s
NVIDIA A100 40GB SXM4
65.33 tok/s
NVIDIA A10G
41.72 tok/s
NVIDIA L4
27.87 tok/s
NVIDIA T4
16.86 tok/s

Efficiency: tok/s per 100W drawn

NVIDIA RTX PRO 6000 Blackwell Workstation Edition
55.04 tok/s / 100W
NVIDIA L4
42.68 tok/s / 100W
NVIDIA H100 80GB HBM3
40 tok/s / 100W
NVIDIA H200
38.11 tok/s / 100W
NVIDIA A100 40GB SXM4
35.03 tok/s / 100W
NVIDIA L40S
34.18 tok/s / 100W
NVIDIA B300
34.13 tok/s / 100W
NVIDIA A10G
33.01 tok/s / 100W
NVIDIA A100 80GB SXM4
28.52 tok/s / 100W
NVIDIA B200
27.85 tok/s / 100W
NVIDIA T4
26.18 tok/s / 100W

Power is the average pulled during the run, sampled at 1Hz. The fastest card is often not the one here, and for anything left running this is the number that shows up on the bill.

Value: tok/s per $1,000 of MSRP

NVIDIA RTX PRO 6000 Blackwell Workstation Edition
17.49 tok/s / $1k
NVIDIA A10G
14.9 tok/s / $1k
NVIDIA L40S
11.45 tok/s / $1k
NVIDIA L4
11.15 tok/s / $1k
NVIDIA T4
7.33 tok/s / $1k
NVIDIA A100 40GB SXM4
5.44 tok/s / $1k
NVIDIA A100 80GB SXM4
3.95 tok/s / $1k
NVIDIA H100 80GB HBM3
3.95 tok/s / $1k
NVIDIA H200
3.86 tok/s / $1k
NVIDIA B300
3.62 tok/s / $1k
NVIDIA B200
3.01 tok/s / $1k

Launch price, not street price, so it ages. A speed leaderboard always crowns the most expensive card; this is the counterweight.

DeepSeek-R1 Distill 14B (Q3_K_M). Measured tokens per second by GPU

NVIDIA RTX PRO 6000 Blackwell Workstation Edition149.8
NVIDIA B300144.7
NVIDIA B200120.5
NVIDIA H200119.7
NVIDIA H100 80GB HBM3118.4
NVIDIA L40S85.85
NVIDIA A100 80GB SXM467.13
NVIDIA A100 40GB SXM465.33
NVIDIA A10G41.72
NVIDIA L427.87
NVIDIA T416.86
GPUtok/sPrompt t/stok/WAvg power
NVIDIA RTX PRO 6000 Blackwell Workstation Edition149.86614.50.55272.2 W
NVIDIA B300144.72457.50.34424.1 W
NVIDIA B200120.55087.30.28432.6 W
NVIDIA H200119.74166.20.38314.0 W
NVIDIA H100 80GB HBM3118.44240.80.4296.1 W
NVIDIA L40S85.854853.10.34251.2 W
NVIDIA A100 80GB SXM467.132010.90.29235.4 W
NVIDIA A100 40GB SXM465.3319800.35186.5 W
NVIDIA A10G41.721706.90.33126.4 W
NVIDIA L427.871457.40.4365.3 W
NVIDIA T416.86597.10.2664.4 W

What the numbers show. Across 11 GPUs measured on our own bench, RTX PRO 6000 Blackwell Workstation Edition is fastest at 150 tok/s. The slowest, T4, manages 16.9, so the spread is 8.9x from top to bottom. The fastest card with 16GB or less is T4 at 16.9 tok/s.

How it compares. H100 80GB HBM3: DeepSeek-R1 Distill 14B (Q3_K_M) 118.4 tok/s, Gemma 3 12B (Q3_K_M) 125.9, Codestral 22B 110.0, Dolphin Mistral 24B Venice 109.1, Dolphin 3.0 R1 Mistral 24B 107.8. 1 of 4 beat DeepSeek-R1 Distill 14B (Q3_K_M) here.

Cost on a rented GPU. 1M generated tokens of DeepSeek-R1 Distill 14B (Q3_K_M): $2.00 on a RTX PRO 6000 Blackwell Workstation Edition ($1.08/hr, 111 min).

DeepSeek-R1 Distill 14B (Q3_K_M): cost per 1M generated tokens on rented GPUs

NVIDIA RTX PRO 6000 Blackwell Workstation Edition$1.08/hr
NVIDIA A100 40GB SXM4$0.47/hr
NVIDIA T4$0.14/hr
NVIDIA L40S$0.79/hr
NVIDIA A100 80GB SXM4$0.95/hr
NVIDIA L4$0.44/hr
NVIDIA H100 80GB HBM3$2.14/hr
NVIDIA H200$3.59/hr
NVIDIA B300$6.94/hr
NVIDIA B200$5.98/hr
GPUCheapest rateSpeed (tok/s)Cost per 1M generated tokens
NVIDIA RTX PRO 6000 Blackwell Workstation Edition$1.08/hr149.8$2.00
NVIDIA A100 40GB SXM4$0.47/hr65.33$2.01
NVIDIA T4$0.14/hr16.86$2.24
NVIDIA L40S$0.79/hr85.85$2.56
NVIDIA A100 80GB SXM4$0.95/hr67.13$3.92
NVIDIA L4$0.44/hr27.87$4.39
NVIDIA H100 80GB HBM3$2.14/hr118.4$5.01
NVIDIA H200$3.59/hr119.7$8.33
NVIDIA B300$6.94/hr144.7$13.32
NVIDIA B200$5.98/hr120.5$13.79

Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.

Speed tiers for DeepSeek-R1 Distill 14B (Q3_K_M). 30+ tok/s: 9 (RTX PRO 6000 Blackwell Workstation Edition, B300, B200); 10-30 tok/s: 2 (L4, T4). 30 tok/s is roughly where replies outpace reading.

Reading your prompt. Before DeepSeek-R1 Distill 14B (Q3_K_M) writes anything it reads the input: 6614.5 tok/s on the RTX PRO 6000 Blackwell Workstation Edition (0.6s for a 4,000-token prompt), 597.1 on the T4 (6.7s). Long documents and big code files feel this number more than the generation speed.

VRAM for DeepSeek-R1 Distill 14B (Q3_K_M). Measured peak 7.5GB, so 8GB is the smallest common card size; smallest card it ran on: T4 (16GB).

Power on DeepSeek-R1 Distill 14B (Q3_K_M). Most efficient: RTX PRO 6000 Blackwell Workstation Edition, 272W, 0.50 kWh per 1M generated tokens. Hungriest: B200, 433W, 1.00 kWh. At $0.15/kWh: $0.076 per 1M generated tokens.

Our verdict

Fastest on DeepSeek-R1 Distill 14B (Q3_K_M): NVIDIA RTX PRO 6000 Blackwell Workstation Edition, 149.8 tok/s. Cheapest to rent per job: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, $2.00 per 1M generated tokens.

FAQ

What GPU do I need to run DeepSeek-R1 Distill 14B (Q3_K_M)?
About 8GB. Smallest card that ran it: NVIDIA T4 (16GB).
How much does it cost to run DeepSeek-R1 Distill 14B (Q3_K_M) in the cloud?
$2.00 per 1M generated tokens on a NVIDIA RTX PRO 6000 Blackwell Workstation Edition at $1.08/hr, cheapest of 10 rentable cards we measured.
Can I run DeepSeek-R1 Distill 14B (Q3_K_M) on a 12GB, 16GB or 24GB card?
It used 7.5GB at the precision we tested. 12GB: yes; 16GB: yes; 24GB: yes.
Is the H100 80GB HBM3 or the A100 80GB SXM4 faster for DeepSeek-R1 Distill 14B (Q3_K_M)?
The H100 80GB HBM3: 118.4 vs 67.13 tok/s, 76% faster on our bench.

How we test

llama.cpp llama-bench at Q3_K_M, 512-token prompt and 128 generated tokens, three runs after a warmup, full GPU offload, with power and VRAM sampled throughout. Token generation is memory-bandwidth-bound, so the ranking tracks bandwidth closely, which makes it a fair guide to cards we haven't run yet.