Gemma 3 12B · 11 GPUs measured first-party · llama.cpp Q4_K_M · Updated October 2026

What GPU Do You Need for Gemma 3 12B?

Gemma 3 12B is the practical middle of Google's open lineup, the size where our '12-14B minimum for serious local work' rule starts being satisfied. Measured on 11 GPUs (llama.cpp, Q4_K_M): 161 tok/s on the B300, ~9GB peak VRAM, so 12GB cards host it with room to spare.

Benchmarked weights: bartowski/google_gemma-3-12b-it-GGUF

Fastest we measured
NVIDIA B300

NVIDIA B300

161.3 tok/s on Gemma 3 12B, the ceiling. Measured on our bench. 288GB of VRAM, $40,000 at launch.

Pros
  • 161.3 tok/s on Gemma 3 12B
  • 288GB, clears the Gemma 3 12B floor
  • Rentable by the hour rather than bought
Cons
  • 1400W board rating
  • Datacenter or workstation hardware, not a retail purchase
Cheapest card that runs it
NVIDIA RTX PRO 6000 Blackwell Workstation Edition

NVIDIA RTX PRO 6000 Blackwell Workstation Edition

138.1 tok/s on Gemma 3 12B, lowest launch price that still fits. Measured on our bench. 96GB of VRAM, $8,565 at launch.

Pros
  • 138.1 tok/s on Gemma 3 12B
  • 96GB, clears the Gemma 3 12B floor
  • Rentable by the hour rather than bought
Cons
  • 600W board rating
  • Datacenter or workstation hardware, not a retail purchase
161.3tok/s
Fastest: NVIDIA B300
measured
11
Cards that run Gemma 3 12B
of 11 we have data for
0
Cards that can't run it at all
published as hard gates, not omissions
598%
Fastest vs slowest that fits
161.3 vs 23.12 tok/s

What GPU Do You Need for Gemma 3 12B?, tok/s by GPU

NVIDIA B300
161.3 tok/s
NVIDIA H200
153.3 tok/s
NVIDIA B200
151.4 tok/s
NVIDIA H100 80GB HBM3
150.6 tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
138.1 tok/s
NVIDIA A100 80GB SXM4
91.45 tok/s
NVIDIA A100 40GB SXM4
88.97 tok/s
NVIDIA L40S
80.92 tok/s
NVIDIA A10G
52.07 tok/s
NVIDIA L4
31.01 tok/s
NVIDIA T4
23.12 tok/s

Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.

Efficiency: tok/s per 100W drawn

NVIDIA H200
101.36 tok/s / 100W
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
73.26 tok/s / 100W
NVIDIA H100 80GB HBM3
64.49 tok/s / 100W
NVIDIA A100 80GB SXM4
57.41 tok/s / 100W
NVIDIA B200
52.08 tok/s / 100W
NVIDIA B300
51.2 tok/s / 100W
NVIDIA A100 40GB SXM4
50.72 tok/s / 100W
NVIDIA L4
47.85 tok/s / 100W
NVIDIA A10G
40.84 tok/s / 100W
NVIDIA L40S
37.02 tok/s / 100W
NVIDIA T4
36.82 tok/s / 100W

Power is the average pulled during the run, sampled at 1Hz. The fastest card is often not the one here, and for anything left running this is the number that shows up on the bill.

Value: tok/s per $1,000 of MSRP

NVIDIA A10G
18.6 tok/s / $1k
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
16.12 tok/s / $1k
NVIDIA L4
12.4 tok/s / $1k
NVIDIA L40S
10.79 tok/s / $1k
NVIDIA T4
10.06 tok/s / $1k
NVIDIA A100 40GB SXM4
7.41 tok/s / $1k
NVIDIA A100 80GB SXM4
5.38 tok/s / $1k
NVIDIA H100 80GB HBM3
5.02 tok/s / $1k
NVIDIA H200
4.94 tok/s / $1k
NVIDIA B300
4.03 tok/s / $1k
NVIDIA B200
3.79 tok/s / $1k

Launch price, not street price, so it ages. A speed leaderboard always crowns the most expensive card; this is the counterweight.

Gemma 3 12B. Measured generation speed by GPU

NVIDIA B300161.3
NVIDIA H200153.3
NVIDIA B200151.4
NVIDIA H100 80GB HBM3150.6
NVIDIA RTX PRO 6000 Blackwell Workstation Edition138.1
NVIDIA A100 80GB SXM491.45
NVIDIA A100 40GB SXM488.97
NVIDIA L40S80.92
NVIDIA A10G52.07
NVIDIA L431.01
NVIDIA T423.12
GPUtok/sPrompt t/stok/WAvg power
NVIDIA B300161.333230.51315.0 W
NVIDIA H200153.353821.01151.2 W
NVIDIA B200151.45768.40.52290.7 W
NVIDIA H100 80GB HBM3150.65432.30.64233.5 W
NVIDIA RTX PRO 6000 Blackwell Workstation Edition138.17537.10.73188.5 W
NVIDIA A100 80GB SXM491.452845.70.57159.3 W
NVIDIA A100 40GB SXM488.972776.30.51175.4 W
NVIDIA L40S80.925987.20.37218.6 W
NVIDIA A10G52.0720600.41127.5 W
NVIDIA L431.011833.10.4864.8 W
NVIDIA T423.12773.90.3762.8 W

A capable middle-of-the-pack citizen. The family verdict applies at 12B too: Gemma is good at everything and the best at nothing. Against its direct rivals in our database, Qwen3 14B (165 tok/s, ~10GB), Phi-4 14B (177 tok/s, ~10GB), DeepSeek-R1 14B (156 tok/s, reasoning), it trades blows without winning a category outright. What it does bring: vision support (the others are text-only), strong multilingual output, and the lightest floor of the group at ~9GB. If your 12GB card needs one do-everything model that can also read screenshots, that combination is a legitimate reason to choose it.

The measurements. 161 tok/s peak, and a notably strong H200 row: 153 tok/s at 151W (1.01 tok/W). At the budget end the L4 manages 31 tok/s, single-user viable, no more. The ~9GB floor is the practical headline: it's the roomiest fit of any 12-14B-class model we benchmark, leaving real context headroom on 12GB cards where its rivals run tight.

How it compares. H100 80GB HBM3: Gemma 3 12B 150.6 tok/s, Qwen3 14B 151.2 (15B), Hermes-4-14B 152.0, Qwen3-14B 152.2, Gemma 4 12B 148.2. 3 of 4 beat Gemma 3 12B here.

Cost on a rented GPU. 1M generated tokens of Gemma 3 12B: $1.47 on a A100 40GB SXM4 ($0.47/hr, 3.1 hours), $11.95 on a B300 ($6.94/hr, 103 min, 8.1x the cost).

Gemma 3 12B: cost per 1M generated tokens on rented GPUs

NVIDIA A100 40GB SXM4$0.47/hr
NVIDIA T4$0.14/hr
NVIDIA RTX PRO 6000 Blackwell Workstation Edition$1.08/hr
NVIDIA L40S$0.79/hr
NVIDIA A100 80GB SXM4$0.95/hr
NVIDIA H100 80GB HBM3$2.14/hr
NVIDIA L4$0.44/hr
NVIDIA H200$3.59/hr
NVIDIA B200$5.98/hr
NVIDIA B300$6.94/hr
GPUCheapest rateSpeed (tok/s)Cost per 1M generated tokens
NVIDIA A100 40GB SXM4$0.47/hr88.97$1.47
NVIDIA T4$0.14/hr23.12$1.63
NVIDIA RTX PRO 6000 Blackwell Workstation Edition$1.08/hr138.1$2.16
NVIDIA L40S$0.79/hr80.92$2.71
NVIDIA A100 80GB SXM4$0.95/hr91.45$2.88
NVIDIA H100 80GB HBM3$2.14/hr150.6$3.94
NVIDIA L4$0.44/hr31.01$3.94
NVIDIA H200$3.59/hr153.3$6.51
NVIDIA B200$5.98/hr151.4$10.97
NVIDIA B300$6.94/hr161.3$11.95

Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.

Speed tiers for Gemma 3 12B. 30+ tok/s: 10 (B300, H200, B200); 10-30 tok/s: 1 (T4). 30 tok/s is roughly where replies outpace reading.

Reading your prompt. Before Gemma 3 12B writes anything it reads the input: 7537.1 tok/s on the RTX PRO 6000 Blackwell Workstation Edition (0.5s for a 4,000-token prompt), 773.9 on the T4 (5.2s). Long documents and big code files feel this number more than the generation speed.

VRAM for Gemma 3 12B. Measured peak 8.1GB, so 12GB is the smallest common card size; smallest card it ran on: T4 (16GB). With long context: Q4_K_M 9GB (tested), Q2_K 6GB, Q3_K_M 8GB, Q5_K_M 11GB, Q6_K 12GB.

Power on Gemma 3 12B. Most efficient: H200, 151W, 0.27 kWh per 1M generated tokens. Hungriest: B300, 315W, 0.54 kWh. At $0.15/kWh: $0.041 per 1M generated tokens.

Our verdict

Gemma 3 12B: 161 tok/s peak and the lightest floor in its class (~9GB), the generalist option at the size where local LLM work starts being real. Choose it for vision + multilingual breadth on a 12GB card; choose its rivals when you need their specializations.

FAQ

What GPU does Gemma 3 12B need?
A 12GB card, comfortably. Measured peak was ~9GB at Q4_K_M, the roomiest fit of any 12-14B model we've benchmarked. Arc B580 ($179) or RTX 3060 12GB both work with context headroom.
Gemma 3 12B or Qwen3 14B / Phi-4 14B?
Measured speeds are close (161 / 165 / 177 tok/s peaks). Gemma wins on vision support and VRAM headroom; Phi on raw speed and polish; Qwen on ecosystem. None dominates, match the temperament to your task.
Can it look at images?
Yes, like the 4B and 27B, Gemma 3 12B is multimodal. If your workflow mixes text tasks with reading screenshots or photos, it's the only 12GB-class option in our lineup that does both.
Is 12B enough for serious local work?
It's the threshold, our working rule is 12-14B minimum for dependable output, with the 27-32B range as the sweet spot. This model sits exactly at that entry line.
What's the best value card for it?
The Arc B580 at $179 is the cheapest comfortable host. On the rental side, the H200's 153 tok/s at 151W was the efficiency standout in our runs.