DeepSeek-R1 Distill 14B · 10 GPUs measured first-party · llama.cpp Q4_K_M · Updated July 2026
DeepSeek-R1 Distill 14B is the mid-size reasoning workhorse, big enough that the thinking traces genuinely sharpen answers, small enough to fit a 12GB card. We measured it on 10 GPUs (llama.cpp, Q4_K_M): 156 tok/s on the B300, ~10GB peak VRAM.
Benchmarked weights: bartowski/DeepSeek-R1-Distill-Qwen-14B-GGUF
What GPU Do You Need for DeepSeek-R1 Distill 14B?, tok/s, fastest 10
Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.
DeepSeek-R1 Distill 14B. Measured generation speed by GPU
| GPU | tok/s | Prompt t/s | tok/W | Avg power |
|---|---|---|---|---|
| NVIDIA B300 | 156.33 | 2964.2 | 0.48 | 328.2 W |
| NVIDIA B200 | 150.69 | 5458.1 | 0.46 | 330.8 W |
| NVIDIA H200 | 146.82 | 4674.7 | 0.72 | 203.1 W |
| NVIDIA H100 80GB HBM3 | 144.82 | 4829.4 | 0.57 | 251.9 W |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 135.41 | 7058.3 | 0.64 | 211.7 W |
| NVIDIA A100 80GB SXM4 | 87 | 2552.9 | 0.58 | 148.9 W |
| NVIDIA A100 40GB SXM4 | 86.76 | 2548.8 | 0.59 | 146.3 W |
| NVIDIA L40S | 74.34 | 5224.4 | 0.44 | 169.1 W |
| NVIDIA A10G | 50.79 | 2148 | 0.41 | 122.9 W |
| NVIDIA L4 | 27.23 | 1527.4 | 0.48 | 56.5 W |
The consensus-tier member. Our thesis for local AI is that a committee of mid-size specialists beats one big generalist, and this model is a founding member of that committee. At 14B the R1 reasoning is real: it catches logic errors the 7B waves through. Run it against Qwen3 14B on the same prompt and you have a two-model panel that fits sequentially on one 12GB card; the answers agreeing is a confidence signal, and their disagreeing is a flag worth reading both traces for.
Hardware math. The ~10GB floor makes a $179 Arc B580 or an RTX 3060 12GB the cheapest hosts, though reasoning models punish slow cards twice, the L4's 27 tok/s means a thousand-token thinking trace takes over half a minute before the answer even starts. If you're running this daily, the A10G tier (51 tok/s) is the floor we'd actually recommend, and the H200's 147 tok/s at 203W is the sensible rented option.
DeepSeek-R1 Distill 14B: 156 tok/s peak, ~10GB floor: the smallest distill whose reasoning we'd call trustworthy, and a natural panel-mate for Qwen3 14B in a consensus setup on a 12GB card. Remember the trace tax: slow cards make reasoning models feel twice as slow.