Llama 3.2 3B · 11 GPUs measured first-party · llama.cpp Q4_K_M · Updated July 2026
Llama 3.2 3B is the bigger sibling in Meta's small-model pair, a solid base at 436 tok/s peak (RTX PRO 6000 Blackwell) with a ~3GB measured floor. We ran it on 11 GPUs with the same pinned llama.cpp Q4_K_M harness as everything else on this site.
Benchmarked weights: bartowski/Llama-3.2-3B-Instruct-GGUF
What GPU Do You Need for Llama 3.2 3B?, tok/s, fastest 11
Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.
Llama 3.2 3B. Measured generation speed by GPU
| GPU | tok/s | Prompt t/s | tok/W | Avg power |
|---|---|---|---|---|
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 435.63 | 20710.1 | 2.61 | 166.7 W |
| NVIDIA B300 | 432.24 | 11154 | 1.62 | 267.5 W |
| NVIDIA H200 | 424.61 | 16433.1 | 2.45 | 173.1 W |
| NVIDIA H100 80GB HBM3 | 422.87 | 16931.6 | 2.18 | 194.4 W |
| NVIDIA B200 | 422.09 | 17563 | 1.3 | 325.1 W |
| NVIDIA L40S | 271.95 | 19089.3 | 1.82 | 149.7 W |
| NVIDIA A100 80GB SXM4 | 264.25 | 8316 | 2.01 | 131.4 W |
| NVIDIA A100 40GB SXM4 | 261.95 | 7916.6 | 2.32 | 113.0 W |
| NVIDIA A10G | 185.5 | 7458 | 1.61 | 115.2 W |
| NVIDIA L4 | 106.62 | 6419.6 | 2.15 | 49.6 W |
| NVIDIA T4 | 85.37 | 2599.5 | 1.53 | 55.8 W |
Where it stands in the small tier. This is a good base model: clean behavior, the full weight of the Llama ecosystem, and enough capability for real summarization and structured-output work. But it sits in the most contested weight class we benchmark, and our honest verdict is that Phi-4 Mini (3.8B, 398 tok/s) answers better for essentially the same hardware budget. Where the 3B wins instead: it's ~10% faster, its ~3GB floor is a touch lighter, and if your deployment story involves fine-tuning, Llama's tooling has no equal.
The measurements. The chart top is flat again, 436, 432, 425 tok/s across three very different cards, because 3B parameters can't stress modern silicon. The practical rows are lower down: 107 tok/s on an L4 at 50W measured, 85 on a T4. Any GPU made in the last five years turns this model into an instant-response tool; the hardware decision is purely about watts and cost.
Llama 3.2 3B: 436 tok/s peak, ~3GB floor, a solid, fast base with unbeatable tooling. As a shipped assistant we'd take Phi-4 Mini's answer quality; as a foundation to tune, or a speed-first workhorse, the Llama is the right pick.