Qwen3 14B · 10 GPUs measured first-party · llama.cpp Q4_K_M · Updated July 2026
Qwen3 14B is the middle child of the lineup: noticeably smarter than the 8B, but in our view not yet the tier where local work gets genuinely good, that starts at 30B-A3B and above. We measured it on 10 GPUs (llama.cpp, Q4_K_M): 165 tok/s on the B300, ~10GB peak VRAM, which makes 12GB the realistic minimum card class.
Benchmarked weights: Qwen/Qwen3-14B-GGUF
What GPU Do You Need for Qwen3 14B?, tok/s, fastest 10
Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.
Qwen3 14B. Measured generation speed by GPU
| GPU | tok/s | Prompt t/s | tok/W | Avg power |
|---|---|---|---|---|
| NVIDIA B300 | 165.44 | 3000.2 | 0.52 | 315.2 W |
| NVIDIA B200 | 158.34 | 5457.1 | 0.5 | 316.1 W |
| NVIDIA H200 | 154.51 | 4882.3 | 1.26 | 122.8 W |
| NVIDIA H100 80GB HBM3 | 151.16 | 5024.5 | 0.73 | 207.7 W |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 139.01 | 7282.1 | 0.8 | 174.8 W |
| NVIDIA A100 40GB SXM4 | 91.28 | 2559.6 | 0.68 | 134.0 W |
| NVIDIA A100 80GB SXM4 | 90.64 | 2554.1 | 0.56 | 162.8 W |
| NVIDIA L40S | 74.88 | 5473.8 | 0.57 | 131.7 W |
| NVIDIA A10G | 51.21 | 2264.7 | 0.45 | 113.2 W |
| NVIDIA L4 | 27.57 | 1575.9 | 0.49 | 55.8 W |
Honest positioning. With 14B you can start doing real local LLM work. It holds context better and hallucinates less than the 8B. But my honest advice is that if your hardware allows it, step up: the 30B-A3B MoE needs 20GB but generates nearly twice as fast in our measurements (309 vs 165 tok/s at the top of each chart) while being a stronger model. 14B is the right stop when your card has 12GB and nothing more: a $179 Arc B580 or an RTX 3060 12GB fits it, and nothing in the 30B class ever will.
Fit and throughput notes. The ~10GB measured peak makes 12GB cards a real but snug fit. Long contexts will push against the ceiling, so keep expectations modest there. On speed, the usual pattern holds: B300 wins the headline at 165 tok/s and 315W, the H200 does 155 tok/s at 123W, 2.4× the efficiency for a 6% speed loss. At the bottom, the L4 manages 28 tok/s, which is workable for a single user but nothing more.
Qwen3 14B: 165 tok/s peak, ~10GB floor: the best model that fits a 12GB card, and that's exactly how we'd use it. If you have 24GB, skip it for the 30B-A3B, which is both faster and stronger. Every number is first-party llama.cpp measurement with logged power.