Phi-4 14B · 10 GPUs measured first-party · llama.cpp Q4_K_M · Updated July 2026
Phi-4 14B is the full-size version of Microsoft's synthetic-data recipe, and the fastest 14B in our database at 177 tok/s (B300), with a ~10GB measured floor. We benchmarked it on 10 GPUs with the same pinned llama.cpp harness as every model on this site.
Benchmarked weights: bartowski/phi-4-GGUF
What GPU Do You Need for Phi-4 14B?, tok/s, fastest 10
Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.
Phi-4 14B. Measured generation speed by GPU
| GPU | tok/s | Prompt t/s | tok/W | Avg power |
|---|---|---|---|---|
| NVIDIA B300 | 177.41 | 3065.9 | 0.53 | 334.4 W |
| NVIDIA H200 | 168.2 | 5134.7 | 0.83 | 203.2 W |
| NVIDIA H100 80GB HBM3 | 165.17 | 5373.6 | 0.69 | 238.0 W |
| NVIDIA B200 | 163.85 | 5796.3 | 0.45 | 365.4 W |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 142.68 | 7777.4 | 0.65 | 219.7 W |
| NVIDIA A100 40GB SXM4 | 99.27 | 2557.7 | 0.64 | 156.2 W |
| NVIDIA A100 80GB SXM4 | 95.22 | 2609.6 | 0.55 | 172.6 W |
| NVIDIA L40S | 74.76 | 5603.3 | 0.49 | 151.9 W |
| NVIDIA A10G | 51.38 | 2382.7 | 0.41 | 124.7 W |
| NVIDIA L4 | 27.14 | 1561.9 | 0.5 | 54.4 W |
Honest positioning: good, not the pick. Phi-4 14B benchmarks beautifully, it out-runs both DeepSeek-R1 14B (156 tok/s) and Qwen3 14B (165 tok/s) at the same weight. But speed isn't why you choose a mid-size local model, and here's our candid ranking: most people running 10-20GB-class models locally are ultimately doing coding and technical work, and in that lane the DeepSeek distill's reasoning, the Dolphin 24Bs' depth, and Mistral's coder models all earn their seats ahead of it. Phi-4's strength is polished general-knowledge answering, real, but a narrower reason to allocate your VRAM.
Where it does fit. If your workload is genuinely general, explanation, writing, Q&A, the speed advantage is real value: 177 tok/s peak and 168 on the H200 at 203W make it the snappiest 14B experience we've measured, and the ~10GB floor keeps 12GB cards in play. As a fast direct-answer counterweight in a panel of slower reasoning models, it also has a legitimate niche: it answers while the others think.
Phi-4 14B: 177 tok/s peak, the fastest 14B we've measured, with a ~10GB floor for 12GB cards. A polished generalist that we nonetheless rank behind DeepSeek, Dolphin and the Mistral coders for the technical work most local users actually do. Choose it for speed and general answers, not as your coding brain.