Llama 3.1 8B · 39 cards measured first-party · Updated July 2026

How Fast Does Llama 3.1 8B Run on Each GPU?

Llama 3.1 8B is our reference LLM, the one workload nearly every GPU in our fleet can run, which makes it the cleanest bandwidth comparison we have. It's also the model most people actually use day to day, because at ~8GB it fits almost anywhere and it's fast enough to feel like a real assistant.

Benchmarked weights: unsloth/Llama-3.1-8B-Instruct-GGUF

Fastest we measured
NVIDIA H100 NVL

NVIDIA H100 NVL

307.8 tok/s on Llama 3.1 8B. Anchored estimate. 94GB of VRAM, 400W board rating. AI Score 67.0/100 across our full 12-workload suite.

Pros
  • 307.8 tok/s on Llama 3.1 8B
  • 94GB, clears the Llama 3.1 8B floor
  • Rentable by the hour rather than bought
Cons
  • 400W board rating
  • Datacenter or workstation hardware, not a retail purchase

Best for: Llama 3.1 8B work where you want the ceiling gone rather than the cheapest entry.

Runner-up
NVIDIA B300

NVIDIA B300

287.23 tok/s on Llama 3.1 8B. Measured on our bench. 288GB of VRAM, 1400W board rating. AI Score 93.8/100 across our full 12-workload suite.

Pros
  • 287.23 tok/s on Llama 3.1 8B
  • 288GB, clears the Llama 3.1 8B floor
  • Rentable by the hour rather than bought
Cons
  • 1400W board rating
  • Datacenter or workstation hardware, not a retail purchase

Best for: Llama 3.1 8B work where you want the ceiling gone rather than the cheapest entry.

Third
NVIDIA B200

NVIDIA B200

274.41 tok/s on Llama 3.1 8B. Measured on our bench. 192GB of VRAM, 1000W board rating. AI Score 78.0/100 across our full 12-workload suite.

Pros
  • 274.41 tok/s on Llama 3.1 8B
  • 192GB, clears the Llama 3.1 8B floor
  • Rentable by the hour rather than bought
Cons
  • 1000W board rating
  • Datacenter or workstation hardware, not a retail purchase

Best for: Llama 3.1 8B work where you want the ceiling gone rather than the cheapest entry.

307.8tok/s
Fastest: NVIDIA H100 NVL
anchored estimate
61
Cards that run Llama 3.1 8B
of 61 we have data for
0
Cards that can't run it at all
published as hard gates, not omissions
714%
Fastest vs slowest that fits
307.8 vs 43.08 tok/s

This is the purest bandwidth test on the site. At 8B the model fits everywhere, nothing is gated, and nothing is compute-bound, so the leaderboard below is close to a direct ranking of memory bandwidth across 100+ GPUs. If you want to understand why VRAM speed matters more than tensor cores for local LLMs, this is the chart that shows it.

Llama 3.1 8B, the 12 fastest cards we have data for

NVIDIA H100 NVL
307.8 tok/s
NVIDIA B300
287.23 tok/s
NVIDIA B200
274.41 tok/s
NVIDIA GH200 Grace Hopper
273.9 tok/s
NVIDIA H200
268.31 tok/s
NVIDIA GeForce RTX 5090
268.14 tok/s
NVIDIA H100 80GB HBM3
261.83 tok/s
NVIDIA H800 80GB
261.8 tok/s
NVIDIA B100
260.7 tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
258.37 tok/s
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
245.5 tok/s
NVIDIA RTX PRO 6000 Blackwell Server Edition
233.8 tok/s

Single stream, batch size 1. 39 of the 61 cards on this page were measured first-party by us; the rest are anchored estimates against those measurements and are labelled in the table below.

Full Llama 3.1 8B leaderboard, every card that runs it

NVIDIA H100 NVL307.8 tok/s
NVIDIA B300287.23 tok/s
NVIDIA B200274.41 tok/s
NVIDIA GH200 Grace Hopper273.9 tok/s
NVIDIA H200268.31 tok/s
NVIDIA GeForce RTX 5090268.14 tok/s
NVIDIA H100 80GB HBM3261.83 tok/s
NVIDIA H800 80GB261.8 tok/s
NVIDIA B100260.7 tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition258.37 tok/s
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition245.5 tok/s
NVIDIA RTX PRO 6000 Blackwell Server Edition233.8 tok/s
NVIDIA RTX PRO 5000 Blackwell208.3 tok/s
NVIDIA GeForce RTX 4090171.29 tok/s
NVIDIA A800 80GB162.6 tok/s
NVIDIA A100 80GB SXM4162.57 tok/s
NVIDIA A100 80GB PCIe161.75 tok/s
NVIDIA RTX 6000 Ada Generation157.89 tok/s
NVIDIA H100 PCIe156.3 tok/s
GeForce RTX 5070 Ti153.23 tok/s
GeForce RTX 5080150.71 tok/s
NVIDIA RTX PRO 4500 Blackwell146.19 tok/s
NVIDIA GeForce RTX 3090144.9 tok/s
NVIDIA GeForce RTX 3080 Ti144.11 tok/s
NVIDIA L40135.12 tok/s
NVIDIA L40S134.96 tok/s
NVIDIA A100 40GB PCIe130.0 tok/s
NVIDIA RTX A5500127.6 tok/s
NVIDIA GeForce RTX 4080127.09 tok/s
NVIDIA GeForce RTX 3080125.85 tok/s
NVIDIA RTX A6000124.87 tok/s
NVIDIA A100 40GB SXM4124.0 tok/s
AMD Radeon Pro W7900122.5 tok/s
NVIDIA RTX A5000119.86 tok/s
NVIDIA RTX 5880 Ada Generation118.3 tok/s
NVIDIA RTX PRO 4000 Blackwell112.36 tok/s
NVIDIA RTX 5000 Ada Generation103.18 tok/s
NVIDIA RTX A4500100.2 tok/s
NVIDIA GeForce RTX 407093.16 tok/s
NVIDIA A10G86.6 tok/s
NVIDIA GeForce RTX 2080 Super85.5 tok/s
NVIDIA Quadro RTX 800084.68 tok/s
NVIDIA Quadro RTX 6000 (Turing)84.21 tok/s
NVIDIA GeForce RTX 2060 Super81.0 tok/s
NVIDIA GeForce RTX 2070 SUPER81.0 tok/s
NVIDIA GeForce RTX 207081.0 tok/s
NVIDIA GeForce RTX 2080 Founders Edition81.0 tok/s
NVIDIA Quadro RTX 500081.0 tok/s
NVIDIA GeForce RTX 3070 Founders Edition80.95 tok/s
AMD Radeon Pro W680080.3 tok/s
AMD Radeon RX 6900 XT80.3 tok/s
NVIDIA RTX 4500 Ada Generation77.5 tok/s
NVIDIA GeForce RTX 3060 Ti76.93 tok/s
NVIDIA RTX A400075.52 tok/s
NVIDIA RTX 4000 (Ada Generation)66.59 tok/s
NVIDIA GeForce RTX 306065.36 tok/s
NVIDIA GeForce RTX 206061.5 tok/s
NVIDIA GeForce RTX 4060 Ti57.09 tok/s
GeForce RTX 406052.4 tok/s
NVIDIA L450.45 tok/s
NVIDIA RTX 2000 Ada Generation43.08 tok/s
GPUResultVRAMSource
NVIDIA H100 NVL307.8 tok/s94GBEstimated
NVIDIA B300287.23 tok/s288GBMeasured
NVIDIA B200274.41 tok/s192GBMeasured
NVIDIA GH200 Grace Hopper273.9 tok/s141GBEstimated
NVIDIA H200268.31 tok/s141GBMeasured
NVIDIA GeForce RTX 5090268.14 tok/s32GBMeasured
NVIDIA H100 80GB HBM3261.83 tok/s80GBMeasured
NVIDIA H800 80GB261.8 tok/s80GBEstimated
NVIDIA B100260.7 tok/s192GBEstimated
NVIDIA RTX PRO 6000 Blackwell Workstation Edition258.37 tok/s96GBMeasured
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition245.5 tok/s96GBEstimated
NVIDIA RTX PRO 6000 Blackwell Server Edition233.8 tok/s96GBMeasured
NVIDIA RTX PRO 5000 Blackwell208.3 tok/s48GBMeasured
NVIDIA GeForce RTX 4090171.29 tok/s24GBMeasured
NVIDIA A800 80GB162.6 tok/s80GBEstimated
NVIDIA A100 80GB SXM4162.57 tok/s80GBMeasured
NVIDIA A100 80GB PCIe161.75 tok/s80GBMeasured
NVIDIA RTX 6000 Ada Generation157.89 tok/s48GBMeasured
NVIDIA H100 PCIe156.3 tok/s80GBEstimated
GeForce RTX 5070 Ti153.23 tok/s16GBMeasured
GeForce RTX 5080150.71 tok/s16GBMeasured
NVIDIA RTX PRO 4500 Blackwell146.19 tok/s32GBMeasured
NVIDIA GeForce RTX 3090144.9 tok/s24GBMeasured
NVIDIA GeForce RTX 3080 Ti144.11 tok/s12GBMeasured
NVIDIA L40135.12 tok/s48GBMeasured
NVIDIA L40S134.96 tok/s48GBMeasured
NVIDIA A100 40GB PCIe130.0 tok/s40GBEstimated
NVIDIA RTX A5500127.6 tok/s24GBEstimated
NVIDIA GeForce RTX 4080127.09 tok/s16GBMeasured
NVIDIA GeForce RTX 3080125.85 tok/s10GBMeasured
NVIDIA RTX A6000124.87 tok/s48GBMeasured
NVIDIA A100 40GB SXM4124.0 tok/s40GBEstimated
AMD Radeon Pro W7900122.5 tok/s48GBEstimated
NVIDIA RTX A5000119.86 tok/s24GBMeasured
NVIDIA RTX 5880 Ada Generation118.3 tok/s48GBEstimated
NVIDIA RTX PRO 4000 Blackwell112.36 tok/s24GBMeasured
NVIDIA RTX 5000 Ada Generation103.18 tok/s32GBMeasured
NVIDIA RTX A4500100.2 tok/s20GBMeasured
NVIDIA GeForce RTX 407093.16 tok/s12GBMeasured
NVIDIA A10G86.6 tok/s24GBMeasured
NVIDIA GeForce RTX 2080 Super85.5 tok/s8GBEstimated
NVIDIA Quadro RTX 800084.68 tok/s48GBMeasured
NVIDIA Quadro RTX 6000 (Turing)84.21 tok/s24GBMeasured
NVIDIA GeForce RTX 2060 Super81.0 tok/s8GBEstimated
NVIDIA GeForce RTX 2070 SUPER81.0 tok/s8GBEstimated
NVIDIA GeForce RTX 207081.0 tok/s8GBEstimated
NVIDIA GeForce RTX 2080 Founders Edition81.0 tok/s8GBEstimated
NVIDIA Quadro RTX 500081.0 tok/s16GBEstimated
NVIDIA GeForce RTX 3070 Founders Edition80.95 tok/s8GBMeasured
AMD Radeon Pro W680080.3 tok/s32GBEstimated
AMD Radeon RX 6900 XT80.3 tok/s16GBEstimated
NVIDIA RTX 4500 Ada Generation77.5 tok/s24GBEstimated
NVIDIA GeForce RTX 3060 Ti76.93 tok/s8GBMeasured
NVIDIA RTX A400075.52 tok/s16GBMeasured
NVIDIA RTX 4000 (Ada Generation)66.59 tok/s20GBMeasured
NVIDIA GeForce RTX 306065.36 tok/s12GBMeasured
NVIDIA GeForce RTX 206061.5 tok/s6GBEstimated
NVIDIA GeForce RTX 4060 Ti57.09 tok/s16GBMeasured
GeForce RTX 406052.4 tok/s8GBMeasured
NVIDIA L450.45 tok/s24GBMeasured
NVIDIA RTX 2000 Ada Generation43.08 tok/s16GBMeasured

Tap any column to sort. Measured = we rented and ran this card ourselves. Estimated = interpolated against our measured anchors, never blended silently.

Because this workload is bandwidth-bound, the ranking above tracks memory bandwidth far more closely than core counts or price. A card with fewer tensor cores and faster memory will beat a card with the opposite.

That's the reason we run twelve workloads instead of publishing one score. A GPU isn't fast or slow. It's fast at some things and gated out of others, and which of those matters depends entirely on what you're actually going to run.

Our verdict

NVIDIA H100 NVL tops our Llama 3.1 8B leaderboard at 307.8 tok/s (anchored estimate), 714% of the way clear of the slowest card that still fits. This is a bandwidth workload: buy memory speed, not tensor cores.

FAQ

What is the fastest GPU for Llama 3.1 8B?
NVIDIA H100 NVL, at 307.8 tok/s on our bench, an anchored estimate against our measured cards. It carries 94GB of VRAM. Of the 61 cards we have Llama 3.1 8B data for, 61 can run it at all.
How much VRAM do I need for Llama 3.1 8B?
~8GB at Q4_K_M. Almost everything clears it, which is exactly what makes it a good yardstick.
Why does the Llama 3.1 8B ranking look different from your other benchmarks?
Because this workload is bandwidth-bound, the ranking above tracks memory bandwidth far more closely than core counts or price. A card with fewer tensor cores and faster memory will beat a card with the opposite. That's why we publish twelve separate workloads rather than one blended score, the ordering genuinely changes depending on the job.
Are these Llama 3.1 8B numbers measured or estimated?
Both, and every row says which. 39 of the 61 cards here were rented and run by us on the same harness. The remainder are anchored estimates interpolated per workload against those measurements. We never blend the two silently, if a row says Estimated, we have not run that card.
Can I rent a GPU to run Llama 3.1 8B instead of buying one?
Yes, and for the cards at the top of this leaderboard it's the only realistic option, most of them have no retail channel at all. It's also how we got these numbers: we rented the hardware by the hour rather than buying it. That's worth considering before you spend on a card to find out whether it's fast enough.
Why publish cards that can't run Llama 3.1 8B?
Nearly every card in our fleet clears this workload's VRAM floor, so there are no gates on this page. That's unusual, most of our suite excludes a meaningful chunk of the market.

How we test

Every ranking on this page comes from our own benchmark runs, not vendor claims. Cards marked Measured were rented and run by us; cards marked Estimated are interpolated per workload against those measured anchors and are labelled on every row, we never blend the two silently. LLMs run on llama.cpp (llama-bench) at Q4_K_M with -p 512 -n 128. Diffusion and video run on diffusers/ComfyUI at BF16, with SDXL at FP16. Each workload gets a warmup pass plus multiple timed runs (5 for small LLMs, 3 for large models and images, 2 for video); we publish the mean as the result and the minimum as the 1% low. Run-to-run variance is under 0.5%. Telemetry, power, temperature, utilisation, clocks, peak VRAM, is sampled at 1 Hz for the duration of every run. Where a model exceeds a card's VRAM we publish a hard won't-fit result rather than quietly dropping to a smaller quantisation. A card that can't run a model scores zero on it. Silently swapping precision to make a number appear would make every number on this site meaningless. All figures are single-GPU, single-stream, batch-size-1. That is the honest way to measure what one card does for one user, and it is deliberately not how a datacenter serves a model. Vendor and MLPerf figures use large batches across many GPUs and will be far higher. Neither is wrong, they answer different questions. Ours answers 'what will this card do for me'.