Gemma 3 4B · 11 GPUs measured first-party · llama.cpp Q4_K_M · Updated July 2026
Gemma 3 4B is Google's small open model, a polished generalist with vision support in a ~4GB package. We measured it on 11 GPUs (llama.cpp, Q4_K_M): 301 tok/s on the B300, with the RTX PRO 6000 essentially tied at 298 while drawing 45% less power.
Benchmarked weights: bartowski/google_gemma-3-4b-it-GGUF
What GPU Do You Need for Gemma 3 4B?, tok/s, fastest 11
Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.
Gemma 3 4B. Measured generation speed by GPU
| GPU | tok/s | Prompt t/s | tok/W | Avg power |
|---|---|---|---|---|
| NVIDIA B300 | 301.07 | 9330.5 | 1.06 | 283.1 W |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 297.73 | 17046.1 | 1.76 | 169.3 W |
| NVIDIA B200 | 291.82 | 14313.1 | 0.88 | 330.8 W |
| NVIDIA H200 | 291.76 | 12728.4 | 1.55 | 188.2 W |
| NVIDIA H100 80GB HBM3 | 289.61 | 13101.8 | 1.37 | 211.0 W |
| NVIDIA L40S | 199.28 | 15129.4 | 1.28 | 155.9 W |
| NVIDIA A100 80GB SXM4 | 176.94 | 7248.6 | 1.45 | 122.2 W |
| NVIDIA A100 40GB SXM4 | 174.84 | 7185.3 | 1.48 | 118.5 W |
| NVIDIA A10G | 135.3 | 6578.9 | 1.12 | 121.0 W |
| NVIDIA L4 | 80.02 | 4952.5 | 1.62 | 49.5 W |
| NVIDIA T4 | 67.26 | 2407.2 | 1.43 | 46.9 W |
The Gemma pattern, stated plainly. Our view of the Gemma 3 line: good models that aren't specialized at anything, middle of the pack across the board rather than best-in-class at one thing. At 4B that's actually a defensible profile: you get Google's training polish, clean multilingual behavior and image understanding in one small package. But the small tier is brutal company, Phi-4 Mini answers better, Qwen3 4B has the wider ecosystem, so Gemma 3 4B earns its slot mainly when you specifically want its vision capability or Google's instruction style in a tiny footprint.
Bench behavior. 301 tok/s peak with the usual flat top (the workstation PRO 6000 at 298 tok/s and 1.76 tok/W is the efficiency pick), and a ~4GB floor that any 6GB card clears. A T4 still manages 67 tok/s, comfortably interactive. Hardware is a non-issue for this model; the decision is entirely about which small model's temperament fits your task.
Gemma 3 4B: 301 tok/s peak, ~4GB floor, vision included, a capable all-rounder in the most competitive weight class we cover. Nothing about it is bad; nothing about it leads. Pick it for the multimodal support or the Google instruction style, not for benchmarks.