Llama 3.2 1B · 24 GPUs measured first-party · llama.cpp Q4_K_M · Updated October 2026

What GPU Do You Need for Llama 3.2 1B?

Llama 3.2 1B holds the single fastest LLM record in our entire database: 893 tokens per second on the RTX PRO 6000 Blackwell. Measured on 24 GPUs (llama.cpp, Q4_K_M) with a ~2GB floor, it's Meta's smallest model, and the benchmark ceiling for what 'fast' means in local inference.

Benchmarked weights: bartowski/Llama-3.2-1B-Instruct-GGUF

Fastest we measured
NVIDIA GeForce RTX 5090

NVIDIA GeForce RTX 5090

1060.1 tok/s on Llama 3.2 1B, the ceiling. Measured on our bench. 32GB of VRAM, $1,999 at launch.

Pros
  • 1060.1 tok/s on Llama 3.2 1B
  • 32GB, clears the Llama 3.2 1B floor
Cons
  • 575W board rating
Best consumer card
NVIDIA GeForce RTX 4090

NVIDIA GeForce RTX 4090

752.5 tok/s on Llama 3.2 1B, fastest card you can buy at retail. Measured on our bench. 24GB of VRAM, $1,599 at launch.

Pros
  • 752.5 tok/s on Llama 3.2 1B
  • 24GB, clears the Llama 3.2 1B floor
Cons
  • 450W board rating
Cheapest card that runs it
NVIDIA GeForce GTX 1660 Super

NVIDIA GeForce GTX 1660 Super

193.4 tok/s on Llama 3.2 1B, lowest launch price that still fits. Measured on our bench. 6GB of VRAM, $229 at launch.

Pros
  • 193.4 tok/s on Llama 3.2 1B
  • 6GB, clears the Llama 3.2 1B floor
Cons
  • 125W board rating
Best value
NVIDIA GeForce RTX 5060

NVIDIA GeForce RTX 5060

377.0 tok/s on Llama 3.2 1B, most speed per dollar. Measured on our bench. 8GB of VRAM, $249 at launch. That is 1513.9 tok/s per $1,000 of launch price.

Pros
  • 377.0 tok/s on Llama 3.2 1B
  • 8GB, clears the Llama 3.2 1B floor
Cons
  • 145W board rating
1060.1tok/s
Fastest: NVIDIA GeForce RTX 5090
measured
11
Cards that run Llama 3.2 1B
of 11 we have data for
0
Cards that can't run it at all
published as hard gates, not omissions
330%
Fastest vs slowest that fits
892.7 vs 207.7 tok/s

What GPU Do You Need for Llama 3.2 1B?, tok/s by GPU

NVIDIA GeForce RTX 5090
1060.1 tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
892.7 tok/s
NVIDIA B300
889.2 tok/s
NVIDIA B200
881.8 tok/s
NVIDIA H100 80GB HBM3
880.6 tok/s
NVIDIA H200
875.5 tok/s
NVIDIA GeForce RTX 4090
752.5 tok/s
GeForce RTX 5080
726.6 tok/s
GeForce RTX 5070 Ti
691.1 tok/s
NVIDIA GeForce RTX 3090
632.2 tok/s
NVIDIA L40S
618.9 tok/s
NVIDIA GeForce RTX 4080
602.3 tok/s
NVIDIA A100 40GB SXM4
532.1 tok/s
NVIDIA A100 80GB SXM4
525.3 tok/s
NVIDIA A10G
400.6 tok/s

Top 15 shown; 9 more cards in the full table below.

Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.

Efficiency: tok/s per 100W drawn

GeForce RTX 5070 Ti
889.42 tok/s / 100W
NVIDIA GeForce RTX 5090
889.38 tok/s / 100W
GeForce RTX 5080
824.71 tok/s / 100W
NVIDIA GeForce RTX 4090
792.96 tok/s / 100W
NVIDIA GeForce RTX 4080
681.31 tok/s / 100W
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
637.67 tok/s / 100W
GeForce RTX 5060 Ti
616.66 tok/s / 100W
NVIDIA H200
607.16 tok/s / 100W
NVIDIA L4
546.27 tok/s / 100W
NVIDIA A100 80GB SXM4
531.16 tok/s / 100W
NVIDIA L40S
528.5 tok/s / 100W
NVIDIA GeForce RTX 5060
525.03 tok/s / 100W
NVIDIA A100 40GB SXM4
472.12 tok/s / 100W
NVIDIA H100 80GB HBM3
467.68 tok/s / 100W
NVIDIA T4
431.83 tok/s / 100W

Top 15 shown; 9 more cards in the full table below.

Power is the average pulled during the run, sampled at 1Hz. The fastest card is often not the one here, and for anything left running this is the number that shows up on the bill.

Value: tok/s per $1,000 of MSRP

NVIDIA GeForce RTX 5060
1513.94 tok/s / $1k
GeForce RTX 5060 Ti
932.89 tok/s / $1k
NVIDIA GeForce RTX 3060
923.1 tok/s / $1k
GeForce RTX 5070 Ti
922.67 tok/s / $1k
NVIDIA GeForce GTX 1660 Super
844.76 tok/s / $1k
NVIDIA GeForce RTX 2060 Super
774.79 tok/s / $1k
GeForce RTX 5080
727.3 tok/s / $1k
NVIDIA GeForce RTX 2070 SUPER
695.83 tok/s / $1k
NVIDIA GeForce RTX 4060 Ti 16GB
567.78 tok/s / $1k
NVIDIA GeForce RTX 5090
530.34 tok/s / $1k
NVIDIA GeForce RTX 4080
502.32 tok/s / $1k
NVIDIA GeForce RTX 4090
470.62 tok/s / $1k
NVIDIA GeForce RTX 3090
421.75 tok/s / $1k
NVIDIA A10G
143.07 tok/s / $1k
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
104.23 tok/s / $1k

Top 15 shown; 9 more cards in the full table below.

Launch price, not street price, so it ages. A speed leaderboard always crowns the most expensive card; this is the counterweight.

Llama 3.2 1B. Measured generation speed by GPU

NVIDIA GeForce RTX 50901060.1
NVIDIA RTX PRO 6000 Blackwell Workstation Edition892.7
NVIDIA B300889.2
NVIDIA B200881.8
NVIDIA H100 80GB HBM3880.6
NVIDIA H200875.5
NVIDIA GeForce RTX 4090752.5
GeForce RTX 5080726.6
GeForce RTX 5070 Ti691.1
NVIDIA GeForce RTX 3090632.2
NVIDIA L40S618.9
NVIDIA GeForce RTX 4080602.3
NVIDIA A100 40GB SXM4532.1
NVIDIA A100 80GB SXM4525.3
NVIDIA A10G400.6
GeForce RTX 5060 Ti400.2
NVIDIA GeForce RTX 5060377
NVIDIA GeForce RTX 2070 SUPER347.2
NVIDIA GeForce RTX 2060 Super309.1
NVIDIA GeForce RTX 3060303.7
NVIDIA GeForce RTX 4060 Ti 16GB283.3
NVIDIA L4259.5
NVIDIA T4207.7
NVIDIA GeForce GTX 1660 Super193.4
GPUtok/sPrompt t/stok/WAvg power
NVIDIA GeForce RTX 50901060.153758.88.89119.2 W
NVIDIA RTX PRO 6000 Blackwell Workstation Edition892.741903.76.38140.0 W
NVIDIA B300889.2257243.38263.3 W
NVIDIA B200881.839531.43293.8 W
NVIDIA H100 80GB HBM3880.636957.24.68188.3 W
NVIDIA H200875.535021.86.07144.2 W
NVIDIA GeForce RTX 4090752.545428.97.9394.9 W
GeForce RTX 5080726.635237.68.2588.1 W
GeForce RTX 5070 Ti691.133777.28.8977.7 W
NVIDIA GeForce RTX 3090632.226122.93.9162.1 W
NVIDIA L40S618.939249.75.28117.1 W
NVIDIA GeForce RTX 4080602.338237.26.8188.4 W
NVIDIA A100 40GB SXM4532.115811.54.72112.7 W
NVIDIA A100 80GB SXM4525.314441.45.3198.9 W
NVIDIA A10G400.617280.44.295.4 W
GeForce RTX 5060 Ti400.218540.76.1764.9 W
NVIDIA GeForce RTX 506037716603.85.2571.8 W
NVIDIA GeForce RTX 2070 SUPER347.29256.53.26106.5 W
NVIDIA GeForce RTX 2060 Super309.18691.13.05101.2 W
NVIDIA GeForce RTX 3060303.710861.63.3590.7 W
NVIDIA GeForce RTX 4060 Ti 16GB283.317002.64.1268.7 W
NVIDIA L4259.517481.35.4647.5 W
NVIDIA T4207.76700.14.3248.1 W
NVIDIA GeForce GTX 1660 Super193.4998.42.5875.1 W

A good base, with one caveat. Llama 3.2 1B is exactly what it says: a clean, well-trained base model with the largest tooling ecosystem in open source behind it. As raw material for fine-tunes and as maximum-velocity pipeline glue, it's excellent. Our honest ranking for the tiny tier, though: if the job is being a small *assistant*, Phi-4 Mini (3.8B) answers noticeably better for a modest speed cost; if the job is running client-side on user devices, Qwen3 0.6B's sub-1B size keeps weak hardware viable. This model's lane is speed itself, batch work where 893 tok/s is the entire specification.

What the record run tells us. The top three cards are separated by barely 1%: 893, 889, 882 tok/s. At 1B parameters, nothing can differentiate big silicon, the model simply can't load it. Which makes the sensible deployment the opposite of glamorous: an L4 does 258 tok/s at 46W measured, a seven-year-old T4 does 201, and both are overkill for most pipelines. This is the model class where hardware stops mattering.

About Llama 3.2 1B. Llama 3.2 1B: from meta-llama, 1.2B parameters, on Hugging Face since September 2024, Llama 3.2 Community licence (gated: accept the terms first). 9,073,830 downloads in the last 30 days and 3 community quantizations.

How it compares. H100 80GB HBM3: Llama 3.2 1B 880.6 tok/s, gemma-3-1b 504.2 (1B), Qwen2.5-1.5B 537.2 (2B), Qwen2.5-Coder-1.5B 538.4 (2B), Qwen2-1.5B 536.5 (2B). Llama 3.2 1B beats all 4 here.

Cost on a rented GPU. 1M generated tokens of Llama 3.2 1B: $0.033 on a RTX 3060 ($0.036/hr, 55 min), $0.10 on a RTX 5090 ($0.39/hr, 16 min, 3.1x the cost).

Llama 3.2 1B: cost per 1M generated tokens on rented GPUs

NVIDIA GeForce RTX 3060$0.036/hr
NVIDIA GeForce RTX 3090$0.12/hr
GeForce RTX 5070 Ti$0.15/hr
NVIDIA GeForce RTX 5060$0.090/hr
GeForce RTX 5080$0.21/hr
NVIDIA GeForce RTX 4080$0.20/hr
GeForce RTX 5060 Ti$0.14/hr
NVIDIA GeForce RTX 5090$0.39/hr
NVIDIA GeForce RTX 4090$0.34/hr
NVIDIA T4$0.14/hr
NVIDIA A100 40GB SXM4$0.47/hr
NVIDIA RTX PRO 6000 Blackwell Workstation Edition$1.08/hr
NVIDIA L40S$0.79/hr
NVIDIA L4$0.44/hr
NVIDIA A100 80GB SXM4$0.95/hr
NVIDIA H100 80GB HBM3$2.14/hr
NVIDIA H200$3.59/hr
NVIDIA B200$5.98/hr
NVIDIA B300$6.94/hr
GPUCheapest rateSpeed (tok/s)Cost per 1M generated tokens
NVIDIA GeForce RTX 3060$0.036/hr303.7$0.033
NVIDIA GeForce RTX 3090$0.12/hr632.2$0.054
GeForce RTX 5070 Ti$0.15/hr691.1$0.060
NVIDIA GeForce RTX 5060$0.090/hr377$0.066
GeForce RTX 5080$0.21/hr726.6$0.080
NVIDIA GeForce RTX 4080$0.20/hr602.3$0.093
GeForce RTX 5060 Ti$0.14/hr400.2$0.094
NVIDIA GeForce RTX 5090$0.39/hr1060.1$0.10
NVIDIA GeForce RTX 4090$0.34/hr752.5$0.12
NVIDIA T4$0.14/hr207.7$0.18
NVIDIA A100 40GB SXM4$0.47/hr532.1$0.25
NVIDIA RTX PRO 6000 Blackwell Workstation Edition$1.08/hr892.7$0.33
NVIDIA L40S$0.79/hr618.9$0.35
NVIDIA L4$0.44/hr259.5$0.47
NVIDIA A100 80GB SXM4$0.95/hr525.3$0.50
NVIDIA H100 80GB HBM3$2.14/hr880.6$0.67
NVIDIA H200$3.59/hr875.5$1.14
NVIDIA B200$5.98/hr881.8$1.88
NVIDIA B300$6.94/hr889.2$2.17

Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.

Speed tiers for Llama 3.2 1B. 30+ tok/s: 24 (RTX 5090, RTX 4090, RTX 5080). 30 tok/s is roughly where replies outpace reading.

Reading your prompt. Before Llama 3.2 1B writes anything it reads the input: 53758.8 tok/s on the RTX 5090 (0.1s for a 4,000-token prompt), 45428.9 on the RTX 4090 (0.1s), 998.4 on the GTX 1660 Super (4.0s). Long documents and big code files feel this number more than the generation speed.

VRAM for Llama 3.2 1B. Measured peak 0.9GB, so 8GB is the smallest common card size; smallest card it ran on: GTX 1660 Super (6GB). With long context: Q4_K_M 1GB (tested), Q2_K 2GB, Q3_K_M 2GB, Q5_K_M 2GB, Q6_K 2GB.

Power on Llama 3.2 1B. Most efficient: RTX 5070 Ti, 78W, 31.2 Wh per 1M generated tokens. Hungriest: B200, 294W, 92.6 Wh.

Our verdict

Llama 3.2 1B: 893 tok/s, the fastest LLM result we've ever measured, with a ~2GB floor that runs on anything. A good base model and the definitive speed play; for tiny-tier assistant quality, our pick shifts to Phi-4 Mini.

FAQ

What's the fastest GPU result for Llama 3.2 1B?
893 tok/s on the RTX PRO 6000 Blackwell, the fastest LLM measurement in our whole database. The B300 and B200 land within 1%: at 1B parameters, top-end hardware is indistinguishable.
What hardware does it actually need?
Almost none, ~2GB peak at Q4_K_M. A T4 from 2018 delivers 201 tok/s. For batch pipelines, the cheapest efficient card you can find (L4: 258 tok/s at 46W) is the right answer.
Llama 3.2 1B or Phi-4 Mini as a small assistant?
Phi-4 Mini, in our view, its 3.8B answers are clearly stronger, and it still measured 398 tok/s peak. Choose this 1B when velocity or fine-tuning is the goal, not conversation quality.
Can it run on user devices like Qwen3 0.6B?
It's borderline, at ~1.2B parameters it's past the sub-1B line where weak devices keep up. On decent laptops yes; for the broadest 'runs on anything the user owns' guarantee, the 0.6B remains the safer bet.
What is it best at?
Being a base. The Llama ecosystem, fine-tuning recipes, tooling, deployment paths, is the deepest in open source, and this is its lightest expression. Speculative decoding drafts, edge summarizers, and custom tunes are its natural jobs.