Gemma 4 12B · 10 GPUs measured first-party · llama.cpp Q4_K_M · Updated July 2026
Gemma 4 12B is the newest Gemma generation in our database, Google's refreshed recipe at the practical mid size. Measured on 10 GPUs (llama.cpp, Q4_K_M): 156 tok/s on the B300, ~9GB peak VRAM, and one standout row: 151 tok/s on the H200 at under 100W.
Benchmarked weights: bartowski/google_gemma-4-12b-it-GGUF
What GPU Do You Need for Gemma 4 12B?, tok/s, fastest 10
Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.
Gemma 4 12B. Measured generation speed by GPU
| GPU | tok/s | Prompt t/s | tok/W | Avg power |
|---|---|---|---|---|
| NVIDIA B300 | 155.96 | 3492.7 | 0.5 | 313.9 W |
| NVIDIA H200 | 151.05 | 5481.5 | 1.52 | 99.6 W |
| NVIDIA H100 80GB HBM3 | 148.17 | 5585 | 0.61 | 243.6 W |
| NVIDIA B200 | 146.36 | 5725.9 | 0.44 | 335.6 W |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 137.73 | 7748.9 | 0.69 | 200.6 W |
| NVIDIA A100 80GB SXM4 | 92.43 | 2840.9 | 0.55 | 166.6 W |
| NVIDIA A100 40GB SXM4 | 90.28 | 2810.8 | 0.65 | 138.6 W |
| NVIDIA L40S | 80.44 | 5909.4 | 0.5 | 160.1 W |
| NVIDIA A10G | 55.98 | 2427.9 | 0.46 | 122.6 W |
| NVIDIA L4 | 31.3 | 1735.4 | 0.62 | 50.4 W |
New generation, familiar profile. Our Gemma family verdict, capable everywhere, specialized nowhere, carries into generation 4 on the evidence so far. Speed and floor land almost exactly on Gemma 3 12B (156 vs 161 tok/s, both ~9GB), so the upgrade case is the newer training, not the hardware math. Same slot, same rivals: it's the do-everything option for a 12GB card, competing against Qwen3 14B's ecosystem, Phi-4 14B's speed, and DeepSeek 14B's reasoning without out-benchmarking any of them.
The efficiency headline. One number in this run deserves its own sentence: the H200 delivered 151 tok/s at 99.6W measured, 1.52 tok/W, the best efficiency we've recorded for any 12B-class model. If you're serving a mid-size model sustained and paying for power, that row is the reason to shortlist this model. At the affordable end, the A10G's 56 tok/s and L4's 31 tok/s keep it interactive on modest silicon.
Gemma 4 12B: 156 tok/s peak, ~9GB floor, and the best 12B-class efficiency we've measured (1.52 tok/W on the H200). The newest Gemma keeps the family character, a polished generalist at the mid size, with a power bill argument the rest of its class can't match.