Phi-4 Mini 3.8B · 11 GPUs measured first-party · llama.cpp Q4_K_M · Updated July 2026
Phi-4 Mini is Microsoft's 3.8B distillation of the Phi-4 recipe, and our pick for the best small assistant in the database. Measured on 11 GPUs (llama.cpp, Q4_K_M): 398 tok/s on the B300, ~4GB peak VRAM, with the RTX PRO 6000 delivering the best efficiency at 2.27 tok/W.
Benchmarked weights: bartowski/microsoft_Phi-4-mini-instruct-GGUF
What GPU Do You Need for Phi-4 Mini 3.8B?, tok/s, fastest 11
Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.
Phi-4 Mini 3.8B. Measured generation speed by GPU
| GPU | tok/s | Prompt t/s | tok/W | Avg power |
|---|---|---|---|---|
| NVIDIA B300 | 398.16 | 10261.4 | 1.32 | 300.8 W |
| NVIDIA H200 | 392.38 | 15554.7 | 2.04 | 192.8 W |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 391.04 | 20459.3 | 2.27 | 172.2 W |
| NVIDIA H100 80GB HBM3 | 388.86 | 16246.1 | 1.85 | 210.2 W |
| NVIDIA B200 | 375.74 | 16160.3 | 1.14 | 328.8 W |
| NVIDIA A100 80GB SXM4 | 239.17 | 7803.5 | 1.9 | 126.0 W |
| NVIDIA A100 40GB SXM4 | 238.48 | 7449.3 | 1.84 | 129.6 W |
| NVIDIA L40S | 229.91 | 17917.7 | 1.5 | 153.0 W |
| NVIDIA A10G | 157.5 | 7137.5 | 1.31 | 120.2 W |
| NVIDIA L4 | 87.04 | 5615.6 | 1.59 | 54.9 W |
| NVIDIA T4 | 66.43 | 2420.1 | 1.21 | 55.0 W |
The small-tier quality pick. In the under-4B class, this is the model we'd actually hand to users. Microsoft's synthetic-textbook training gives it answer quality that punches above its parameter count, in our experience clearly ahead of Llama 3.2's small pair for assistant work, which is why our tiny-tier advice reads: Qwen3 0.6B for client-side deployment, Phi-4 Mini for the best small brain. At 398 tok/s peak it gives up only ~10% speed to the Llama 3B while being the noticeably better conversationalist.
Deployment math. The ~4GB floor is the sweet spot for modest hardware: 6GB laptop GPUs and every desktop card from the last half-decade clear it with room. The chart's flat top (398/392/391 across three card classes) says the usual: small models can't use big silicon, so deploy on efficiency: the PRO 6000's 391 tok/s at 172W is elegant, but an L4 at 87 tok/s and 55W is the volume-serving bargain.
Phi-4 Mini: 398 tok/s peak, ~4GB floor, and the best answers-per-parameter in our small-model lineup, our quality pick for the tiny tier. Fits nearly anything, including modest laptop GPUs; pair it with a big-model escalation path and most everyday queries never need the big model.