Phi-4 14B · 11 GPUs measured first-party · llama.cpp Q4_K_M · Updated October 2026
Phi-4 14B is the full-size version of Microsoft's synthetic-data recipe, and the fastest 14B in our database at 177 tok/s (B300), with a ~10GB measured floor. We benchmarked it on 11 GPUs with the same pinned llama.cpp harness as every model on this site.
Benchmarked weights: bartowski/phi-4-GGUF

177.4 tok/s on Phi-4 14B, the ceiling. Measured on our bench. 288GB of VRAM, $40,000 at launch.

142.7 tok/s on Phi-4 14B, lowest launch price that still fits. Measured on our bench. 96GB of VRAM, $8,565 at launch.
What GPU Do You Need for Phi-4 14B?, tok/s by GPU
Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.
Efficiency: tok/s per 100W drawn
Power is the average pulled during the run, sampled at 1Hz. The fastest card is often not the one here, and for anything left running this is the number that shows up on the bill.
Value: tok/s per $1,000 of MSRP
Launch price, not street price, so it ages. A speed leaderboard always crowns the most expensive card; this is the counterweight.
Phi-4 14B. Measured generation speed by GPU
| GPU | tok/s | Prompt t/s | tok/W | Avg power |
|---|---|---|---|---|
| NVIDIA B300 | 177.4 | 3065.9 | 0.53 | 334.4 W |
| NVIDIA H200 | 168.2 | 5134.7 | 0.83 | 203.2 W |
| NVIDIA H100 80GB HBM3 | 165.2 | 5373.6 | 0.69 | 238.0 W |
| NVIDIA B200 | 163.8 | 5796.3 | 0.45 | 365.4 W |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 142.7 | 7777.4 | 0.65 | 219.7 W |
| NVIDIA A100 40GB SXM4 | 98.65 | 2536.6 | 0.49 | 200.1 W |
| NVIDIA A100 80GB SXM4 | 95.22 | 2609.6 | 0.55 | 172.6 W |
| NVIDIA L40S | 74.98 | 5688.1 | 0.31 | 238.9 W |
| NVIDIA A10G | 48.18 | 1815.6 | 0.38 | 128.4 W |
| NVIDIA L4 | 27.24 | 1623.7 | 0.41 | 65.8 W |
| NVIDIA T4 | 17.46 | 662 | 0.28 | 62.9 W |
Honest positioning: good, not the pick. Phi-4 14B benchmarks beautifully, it out-runs both DeepSeek-R1 14B (156 tok/s) and Qwen3 14B (165 tok/s) at the same weight. But speed isn't why you choose a mid-size local model, and here's our candid ranking: most people running 10-20GB-class models locally are ultimately doing coding and technical work, and in that lane the DeepSeek distill's reasoning, the Dolphin 24Bs' depth, and Mistral's coder models all earn their seats ahead of it. Phi-4's strength is polished general-knowledge answering, real, but a narrower reason to allocate your VRAM.
Where it does fit. If your workload is genuinely general, explanation, writing, Q&A, the speed advantage is real value: 177 tok/s peak and 168 on the H200 at 203W make it the snappiest 14B experience we've measured, and the ~10GB floor keeps 12GB cards in play. As a fast direct-answer counterweight in a panel of slower reasoning models, it also has a legitimate niche: it answers while the others think.
How it compares. H100 80GB HBM3: Phi-4 14B 165.2 tok/s, GLM-4.7-Flash-REAP-23B-A3B 168.4, gemma-2-9b 174.7 (9B), NemoMix-Unleashed-12B 177.3, Mistral-Nemo-Instruct-2407 177.4. All 4 beat Phi-4 14B here. GLM-4.7-Flash-REAP-23B-A3B is a mixture-of-experts, so per token it computes only a slice of its size.
Cost on a rented GPU. 1M generated tokens of Phi-4 14B: $1.33 on a A100 40GB SXM4 ($0.47/hr, 2.8 hours), $10.87 on a B300 ($6.94/hr, 94 min, 8.2x the cost).
Phi-4 14B: cost per 1M generated tokens on rented GPUs
| GPU | Cheapest rate | Speed (tok/s) | Cost per 1M generated tokens |
|---|---|---|---|
| NVIDIA A100 40GB SXM4 | $0.47/hr | 98.65 | $1.33 |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | $1.08/hr | 142.7 | $2.09 |
| NVIDIA T4 | $0.14/hr | 17.46 | $2.16 |
| NVIDIA A100 80GB SXM4 | $0.95/hr | 95.22 | $2.76 |
| NVIDIA L40S | $0.79/hr | 74.98 | $2.93 |
| NVIDIA H100 80GB HBM3 | $2.14/hr | 165.2 | $3.59 |
| NVIDIA L4 | $0.44/hr | 27.24 | $4.49 |
| NVIDIA H200 | $3.59/hr | 168.2 | $5.93 |
| NVIDIA B200 | $5.98/hr | 163.8 | $10.14 |
| NVIDIA B300 | $6.94/hr | 177.4 | $10.87 |
Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.
Speed tiers for Phi-4 14B. 30+ tok/s: 9 (B300, H200, H100 80GB HBM3); 10-30 tok/s: 2 (L4, T4). 30 tok/s is roughly where replies outpace reading.
Reading your prompt. Before Phi-4 14B writes anything it reads the input: 7777.4 tok/s on the RTX PRO 6000 Blackwell Workstation Edition (0.5s for a 4,000-token prompt), 662.0 on the T4 (6.0s). Long documents and big code files feel this number more than the generation speed.
VRAM for Phi-4 14B. Measured peak 9.0GB, so 12GB is the smallest common card size; smallest card it ran on: T4 (16GB). With long context: Q4_K_M 11GB (tested), Q2_K 7GB, Q3_K_M 9GB, Q5_K_M 13GB, Q6_K 15GB.
Power on Phi-4 14B. Most efficient: H200, 203W, 0.34 kWh per 1M generated tokens. Hungriest: B200, 365W, 0.62 kWh. At $0.15/kWh: $0.050 per 1M generated tokens.
Phi-4 14B: 177 tok/s peak, the fastest 14B we've measured, with a ~10GB floor for 12GB cards. A polished generalist that we nonetheless rank behind DeepSeek, Dolphin and the Mistral coders for the technical work most local users actually do. Choose it for speed and general answers, not as your coding brain.