Gemma 3 12B · 11 GPUs measured first-party · llama.cpp Q4_K_M · Updated July 2026
Gemma 3 12B is the practical middle of Google's open lineup, the size where our '12-14B minimum for serious local work' rule starts being satisfied. Measured on 11 GPUs (llama.cpp, Q4_K_M): 161 tok/s on the B300, ~9GB peak VRAM, so 12GB cards host it with room to spare.
Benchmarked weights: bartowski/google_gemma-3-12b-it-GGUF
What GPU Do You Need for Gemma 3 12B?, tok/s, fastest 11
Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.
Gemma 3 12B. Measured generation speed by GPU
| GPU | tok/s | Prompt t/s | tok/W | Avg power |
|---|---|---|---|---|
| NVIDIA B300 | 161.27 | 3323 | 0.51 | 315.0 W |
| NVIDIA H200 | 153.26 | 5382 | 1.01 | 151.2 W |
| NVIDIA B200 | 151.41 | 5768.4 | 0.52 | 290.7 W |
| NVIDIA H100 80GB HBM3 | 150.58 | 5432.3 | 0.64 | 233.5 W |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 138.09 | 7537.1 | 0.73 | 188.5 W |
| NVIDIA A100 80GB SXM4 | 91.45 | 2845.7 | 0.57 | 159.3 W |
| NVIDIA A100 40GB SXM4 | 89.94 | 2855.6 | 0.83 | 109.0 W |
| NVIDIA L40S | 80.38 | 5895.3 | 0.47 | 172.2 W |
| NVIDIA A10G | 55.67 | 2483.6 | 0.44 | 125.6 W |
| NVIDIA L4 | 30.99 | 1797 | 0.55 | 56.5 W |
| NVIDIA T4 | 21.85 | 768.2 | 0.38 | 58.2 W |
A capable middle-of-the-pack citizen. The family verdict applies at 12B too: Gemma is good at everything and the best at nothing. Against its direct rivals in our database, Qwen3 14B (165 tok/s, ~10GB), Phi-4 14B (177 tok/s, ~10GB), DeepSeek-R1 14B (156 tok/s, reasoning), it trades blows without winning a category outright. What it does bring: vision support (the others are text-only), strong multilingual output, and the lightest floor of the group at ~9GB. If your 12GB card needs one do-everything model that can also read screenshots, that combination is a legitimate reason to choose it.
The measurements. 161 tok/s peak, and a notably strong H200 row: 153 tok/s at 151W (1.01 tok/W). At the budget end the L4 manages 31 tok/s, single-user viable, no more. The ~9GB floor is the practical headline: it's the roomiest fit of any 12-14B-class model we benchmark, leaving real context headroom on 12GB cards where its rivals run tight.
Gemma 3 12B: 161 tok/s peak and the lightest floor in its class (~9GB), the generalist option at the size where local LLM work starts being real. Choose it for vision + multilingual breadth on a 12GB card; choose its rivals when you need their specializations.