Llama 3.1 8B · 71 cards measured first-party · Updated October 2026

How Fast Does Llama 3.1 8B Run on Each GPU?

Llama 3.1 8B is our reference LLM, the one workload nearly every GPU in our fleet can run, which makes it the cleanest bandwidth comparison we have. It's also the model most people actually use day to day, because at ~8GB it fits almost anywhere and it's fast enough to feel like a real assistant.

Benchmarked weights: unsloth/Llama-3.1-8B-Instruct-GGUF

Fastest we measured
NVIDIA B300

NVIDIA B300

287.2 tok/s on Llama 3.1 8B, the ceiling. Measured on our bench. 288GB of VRAM, $40,000 at launch.

Pros
  • 287.2 tok/s on Llama 3.1 8B
  • 288GB, clears the Llama 3.1 8B floor
  • Rentable by the hour rather than bought
Cons
  • 1400W board rating
  • Datacenter or workstation hardware, not a retail purchase
Best consumer card
NVIDIA GeForce RTX 5090

NVIDIA GeForce RTX 5090

268.1 tok/s on Llama 3.1 8B, fastest card you can buy at retail. Measured on our bench. 32GB of VRAM, $1,999 at launch.

Pros
  • 268.1 tok/s on Llama 3.1 8B
  • 32GB, clears the Llama 3.1 8B floor
Cons
  • 575W board rating
Cheapest card that runs it
Intel Arc B580

Intel Arc B580

60.3 tok/s on Llama 3.1 8B, lowest launch price that still fits. Anchored estimate. 12GB of VRAM, $179 at launch.

Pros
  • 60.3 tok/s on Llama 3.1 8B
  • 12GB, clears the Llama 3.1 8B floor
Cons
  • 190W board rating
Best value
NVIDIA GeForce RTX 5060

NVIDIA GeForce RTX 5060

79.05 tok/s on Llama 3.1 8B, most speed per dollar. Measured on our bench. 8GB of VRAM, $249 at launch. That is 317.5 tok/s per $1,000 of launch price.

Pros
  • 79.05 tok/s on Llama 3.1 8B
  • 8GB, clears the Llama 3.1 8B floor
Cons
  • 145W board rating
287.2tok/s
Fastest: NVIDIA B300
measured
100
Cards that run Llama 3.1 8B
of 102 we have data for
2
Cards that can't run it at all
published as hard gates, not omissions
755%
Fastest vs slowest that fits
287.2 vs 33.6 tok/s

This is the purest bandwidth test on the site. At 8B the model fits everywhere, nothing is gated, and nothing is compute-bound, so the leaderboard below is close to a direct ranking of memory bandwidth across 100+ GPUs. If you want to understand why VRAM speed matters more than tensor cores for local LLMs, this is the chart that shows it.

Llama 3.1 8B: speed on every GPU we have data for

NVIDIA B300
287.2 tok/s
NVIDIA B200
274.4 tok/s
NVIDIA GH200 Grace Hopper
273.9 tok/s
NVIDIA H200
268.3 tok/s
NVIDIA GeForce RTX 5090
268.1 tok/s
NVIDIA H100 80GB HBM3
261.8 tok/s
NVIDIA H800 80GB
261.8 tok/s
NVIDIA B100
260.7 tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
258.4 tok/s
NVIDIA RTX PRO 6000 Blackwell Server Edition
233.8 tok/s
NVIDIA H100 NVL
231.9 tok/s
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
231.5 tok/s
NVIDIA RTX PRO 5000 Blackwell
208.3 tok/s
NVIDIA H100 PCIe
194.3 tok/s
NVIDIA GeForce RTX 4090
171.3 tok/s

Top 15 shown; 85 more cards in the full table below.

Single stream, batch size 1. 39 of the 61 cards on this page were measured first-party by us; the rest are anchored estimates against those measurements and are labelled in the table below.

Efficiency: tok/s per 100W drawn

NVIDIA GeForce RTX 5090
284.65 tok/s / 100W
NVIDIA RTX PRO 4500 Blackwell
147.37 tok/s / 100W
NVIDIA H100 80GB HBM3
129.81 tok/s / 100W
NVIDIA H200
125.91 tok/s / 100W
NVIDIA RTX PRO 5000 Blackwell
122.03 tok/s / 100W
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
121.82 tok/s / 100W
NVIDIA H100 PCIe
113.62 tok/s / 100W
NVIDIA RTX PRO 6000 Blackwell Server Edition
107.64 tok/s / 100W
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
107.07 tok/s / 100W
NVIDIA H100 NVL
106.31 tok/s / 100W
GeForce RTX 5080
104.51 tok/s / 100W
NVIDIA RTX PRO 4000 Blackwell
100.59 tok/s / 100W
NVIDIA A100 80GB PCIe
97.91 tok/s / 100W
GeForce RTX 4080 Super
97.45 tok/s / 100W
NVIDIA RTX 2000 Ada Generation
97.03 tok/s / 100W

Top 15 shown; 54 more cards in the full table below.

Power is the average pulled during the run, sampled at 1Hz. The fastest card is often not the one here, and for anything left running this is the number that shows up on the bill.

Value: tok/s per $1,000 of MSRP

NVIDIA GeForce RTX 5060
317.47 tok/s / $1k
GeForce RTX 5070
218.34 tok/s / $1k
NVIDIA GeForce GTX 1660 Super
212.4 tok/s / $1k
GeForce RTX 5070 Ti
204.58 tok/s / $1k
NVIDIA GeForce RTX 3060
198.66 tok/s / $1k
GeForce RTX 5060 Ti
196.27 tok/s / $1k
NVIDIA GeForce RTX 3060 Ti
192.81 tok/s / $1k
NVIDIA GeForce RTX 3080
180.04 tok/s / $1k
NVIDIA GeForce GTX 1660 Ti
176.67 tok/s / $1k
GeForce RTX 4060
175.25 tok/s / $1k
NVIDIA GeForce RTX 3070 Ti
170.5 tok/s / $1k
NVIDIA GeForce RTX 3070 Founders Edition
162.22 tok/s / $1k
NVIDIA GeForce GTX 1660
156.94 tok/s / $1k
NVIDIA GeForce RTX 4070 Super
156.29 tok/s / $1k
NVIDIA GeForce RTX 4070
155.53 tok/s / $1k

Top 15 shown; 55 more cards in the full table below.

Launch price, not street price, so it ages. A speed leaderboard always crowns the most expensive card; this is the counterweight.

Full Llama 3.1 8B leaderboard, every card that runs it

NVIDIA B300287.2 tok/s
NVIDIA B200274.4 tok/s
NVIDIA GH200 Grace Hopper273.9 tok/s
NVIDIA H200268.3 tok/s
NVIDIA GeForce RTX 5090268.1 tok/s
NVIDIA H100 80GB HBM3261.8 tok/s
NVIDIA H800 80GB261.8 tok/s
NVIDIA B100260.7 tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition258.4 tok/s
NVIDIA RTX PRO 6000 Blackwell Server Edition233.8 tok/s
NVIDIA H100 NVL231.9 tok/s
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition231.5 tok/s
NVIDIA RTX PRO 5000 Blackwell208.3 tok/s
NVIDIA H100 PCIe194.3 tok/s
NVIDIA GeForce RTX 4090171.3 tok/s
NVIDIA A800 80GB162.6 tok/s
NVIDIA A100 80GB SXM4162.6 tok/s
NVIDIA A100 80GB PCIe161.8 tok/s
NVIDIA GeForce RTX 3090 Ti161.0 tok/s
NVIDIA RTX 6000 Ada Generation157.9 tok/s
NVIDIA A100 40GB SXM4157.3 tok/s
NVIDIA RTX 5880 Ada Generation157.2 tok/s
NVIDIA A100 40GB PCIe154.2 tok/s
GeForce RTX 5070 Ti153.2 tok/s
GeForce RTX 5080150.7 tok/s
AMD Radeon RX 7900 XTX147.1 tok/s
NVIDIA RTX PRO 4500 Blackwell146.2 tok/s
NVIDIA GeForce RTX 3090144.9 tok/s
NVIDIA GeForce RTX 3080 Ti144.1 tok/s
NVIDIA L40S135.5 tok/s
NVIDIA L40135.1 tok/s
NVIDIA RTX A5500127.6 tok/s
NVIDIA GeForce RTX 4080127.1 tok/s
GeForce RTX 4080 Super126.3 tok/s
NVIDIA GeForce RTX 3080125.8 tok/s
NVIDIA RTX A6000124.9 tok/s
AMD Radeon Pro W7900122.5 tok/s
GeForce RTX 5070119.9 tok/s
NVIDIA RTX A5000119.9 tok/s
NVIDIA GeForce RTX 4070 Ti Super119.6 tok/s
AMD Radeon RX 7900 XT115.5 tok/s
NVIDIA RTX PRO 4000 Blackwell112.4 tok/s
NVIDIA Titan RTX107.4 tok/s
NVIDIA A40106.2 tok/s
NVIDIA RTX 5000 Ada Generation103.2 tok/s
AMD Radeon RX 9070 XT103.1 tok/s
NVIDIA TITAN V102.3 tok/s
NVIDIA GeForce RTX 3070 Ti102.1 tok/s
NVIDIA RTX A4500100.2 tok/s
AMD Radeon RX 907099.0 tok/s
NVIDIA GeForce RTX 2080 Ti Founders Edition98.34 tok/s
AMD Radeon RX 7800 XT97.2 tok/s
NVIDIA GeForce RTX 4070 Super93.62 tok/s
NVIDIA GeForce RTX 407093.16 tok/s
NVIDIA GeForce RTX 4070 Ti92.09 tok/s
NVIDIA A10G86.62 tok/s
NVIDIA GeForce RTX 2080 Super85.5 tok/s
NVIDIA Quadro RTX 800084.68 tok/s
AMD Radeon RX 6900 XT84.3 tok/s
NVIDIA Quadro RTX 6000 (Turing)84.21 tok/s
GeForce RTX 5060 Ti84.2 tok/s
AMD Radeon Pro W780084.1 tok/s
AMD Radeon RX 6950 XT84.1 tok/s
NVIDIA GeForce RTX 2080 Founders Edition81.0 tok/s
NVIDIA GeForce RTX 3070 Founders Edition80.95 tok/s
AMD Radeon Pro W680080.3 tok/s
AMD Radeon RX 6800 XT80.3 tok/s
AMD Radeon RX 680080.3 tok/s
NVIDIA GeForce RTX 506079.05 tok/s
NVIDIA RTX 4500 Ada Generation78.87 tok/s
NVIDIA GeForce RTX 3060 Ti76.93 tok/s
NVIDIA RTX A400075.52 tok/s
NVIDIA GeForce RTX 2070 SUPER75.3 tok/s
NVIDIA Quadro RTX 500073.61 tok/s
NVIDIA GeForce RTX 2070 (power capped)68.09 tok/s
NVIDIA RTX 4000 (Ada Generation)66.59 tok/s
AMD Radeon RX 7700 XT65.8 tok/s
NVIDIA GeForce RTX 306065.36 tok/s
NVIDIA GeForce RTX 2060 Super (power capped)62.0 tok/s
NVIDIA TITAN X (Pascal)61.6 tok/s
Intel Arc B58060.3 tok/s
NVIDIA GeForce RTX 4060 Ti 16GB57.09 tok/s
GeForce RTX 406052.4 tok/s
GeForce GTX 1080 Ti50.52 tok/s
NVIDIA L450.45 tok/s
NVIDIA GeForce GTX 1660 Ti49.29 tok/s
NVIDIA GeForce GTX 1660 Super48.64 tok/s
AMD Radeon RX 670048.5 tok/s
NVIDIA GeForce RTX 505046.6 tok/s
Intel Arc A770 Limited Edition46.3 tok/s
NVIDIA TITAN Xp (power capped)46.06 tok/s
AMD Radeon RX 760044.1 tok/s
NVIDIA RTX 2000 Ada Generation43.08 tok/s
NVIDIA GeForce RTX 305041.4 tok/s
Intel Arc A75037.0 tok/s
NVIDIA GeForce GTX 108035.98 tok/s
NVIDIA GeForce GTX 1070 Ti35.9 tok/s
NVIDIA T435.01 tok/s
NVIDIA GeForce GTX 166034.37 tok/s
Intel Arc Pro A6033.6 tok/s
GPUResultVRAMSource
NVIDIA B300287.2 tok/s288GBMeasured
NVIDIA B200274.4 tok/s192GBMeasured
NVIDIA GH200 Grace Hopper273.9 tok/s141GBEstimated
NVIDIA H200268.3 tok/s141GBMeasured
NVIDIA GeForce RTX 5090268.1 tok/s32GBMeasured
NVIDIA H100 80GB HBM3261.8 tok/s80GBMeasured
NVIDIA H800 80GB261.8 tok/s80GBEstimated
NVIDIA B100260.7 tok/s192GBEstimated
NVIDIA RTX PRO 6000 Blackwell Workstation Edition258.4 tok/s96GBMeasured
NVIDIA RTX PRO 6000 Blackwell Server Edition233.8 tok/s96GBMeasured
NVIDIA H100 NVL231.9 tok/s94GBMeasured
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition231.5 tok/s96GBMeasured
NVIDIA RTX PRO 5000 Blackwell208.3 tok/s48GBMeasured
NVIDIA H100 PCIe194.3 tok/s80GBMeasured
NVIDIA GeForce RTX 4090171.3 tok/s24GBMeasured
NVIDIA A800 80GB162.6 tok/s80GBEstimated
NVIDIA A100 80GB SXM4162.6 tok/s80GBMeasured
NVIDIA A100 80GB PCIe161.8 tok/s80GBMeasured
NVIDIA GeForce RTX 3090 Ti161.0 tok/s24GBMeasured
NVIDIA RTX 6000 Ada Generation157.9 tok/s48GBMeasured
NVIDIA A100 40GB SXM4157.3 tok/s40GBMeasured
NVIDIA RTX 5880 Ada Generation157.2 tok/s48GBMeasured
NVIDIA A100 40GB PCIe154.2 tok/s40GBMeasured
GeForce RTX 5070 Ti153.2 tok/s16GBMeasured
GeForce RTX 5080150.7 tok/s16GBMeasured
AMD Radeon RX 7900 XTX147.1 tok/s24GBEstimated
NVIDIA RTX PRO 4500 Blackwell146.2 tok/s32GBMeasured
NVIDIA GeForce RTX 3090144.9 tok/s24GBMeasured
NVIDIA GeForce RTX 3080 Ti144.1 tok/s12GBMeasured
NVIDIA L40S135.5 tok/s48GBMeasured
NVIDIA L40135.1 tok/s48GBMeasured
NVIDIA RTX A5500127.6 tok/s24GBEstimated
NVIDIA GeForce RTX 4080127.1 tok/s16GBMeasured
GeForce RTX 4080 Super126.3 tok/s16GBMeasured
NVIDIA GeForce RTX 3080125.8 tok/s10GBMeasured
NVIDIA RTX A6000124.9 tok/s48GBMeasured
AMD Radeon Pro W7900122.5 tok/s48GBEstimated
GeForce RTX 5070119.9 tok/s12GBMeasured
NVIDIA RTX A5000119.9 tok/s24GBMeasured
NVIDIA GeForce RTX 4070 Ti Super119.6 tok/s16GBMeasured
AMD Radeon RX 7900 XT115.5 tok/s20GBEstimated
NVIDIA RTX PRO 4000 Blackwell112.4 tok/s24GBMeasured
NVIDIA Titan RTX107.4 tok/s24GBMeasured
NVIDIA A40106.2 tok/s48GBMeasured
NVIDIA RTX 5000 Ada Generation103.2 tok/s32GBMeasured
AMD Radeon RX 9070 XT103.1 tok/s16GBEstimated
NVIDIA TITAN V102.3 tok/s12GBMeasured
NVIDIA GeForce RTX 3070 Ti102.1 tok/s8GBMeasured
NVIDIA RTX A4500100.2 tok/s20GBMeasured
AMD Radeon RX 907099.0 tok/s16GBEstimated
NVIDIA GeForce RTX 2080 Ti Founders Edition98.34 tok/s11GBMeasured
AMD Radeon RX 7800 XT97.2 tok/s16GBEstimated
NVIDIA GeForce RTX 4070 Super93.62 tok/s12GBMeasured
NVIDIA GeForce RTX 407093.16 tok/s12GBMeasured
NVIDIA GeForce RTX 4070 Ti92.09 tok/s12GBMeasured
NVIDIA A10G86.62 tok/s24GBMeasured
NVIDIA GeForce RTX 2080 Super85.5 tok/s8GBEstimated
NVIDIA Quadro RTX 800084.68 tok/s48GBMeasured
AMD Radeon RX 6900 XT84.3 tok/s16GBEstimated
NVIDIA Quadro RTX 6000 (Turing)84.21 tok/s24GBMeasured
GeForce RTX 5060 Ti84.2 tok/s16GBMeasured
AMD Radeon Pro W780084.1 tok/s32GBEstimated
AMD Radeon RX 6950 XT84.1 tok/s16GBEstimated
NVIDIA GeForce RTX 2080 Founders Edition81.0 tok/s8GBEstimated
NVIDIA GeForce RTX 3070 Founders Edition80.95 tok/s8GBMeasured
AMD Radeon Pro W680080.3 tok/s32GBEstimated
AMD Radeon RX 6800 XT80.3 tok/s16GBEstimated
AMD Radeon RX 680080.3 tok/s16GBEstimated
NVIDIA GeForce RTX 506079.05 tok/s8GBMeasured
NVIDIA RTX 4500 Ada Generation78.87 tok/s24GBMeasured
NVIDIA GeForce RTX 3060 Ti76.93 tok/s8GBMeasured
NVIDIA RTX A400075.52 tok/s16GBMeasured
NVIDIA GeForce RTX 2070 SUPER75.3 tok/s8GBMeasured
NVIDIA Quadro RTX 500073.61 tok/s16GBMeasured
NVIDIA GeForce RTX 2070 (power capped)68.09 tok/s8GBMeasured
NVIDIA RTX 4000 (Ada Generation)66.59 tok/s20GBMeasured
AMD Radeon RX 7700 XT65.8 tok/s12GBEstimated
NVIDIA GeForce RTX 306065.36 tok/s12GBMeasured
NVIDIA GeForce RTX 2060 Super (power capped)62.0 tok/s8GBMeasured
NVIDIA TITAN X (Pascal)61.6 tok/s12GBEstimated
Intel Arc B58060.3 tok/s12GBEstimated
NVIDIA GeForce RTX 4060 Ti 16GB57.09 tok/s16GBMeasured
GeForce RTX 406052.4 tok/s8GBMeasured
GeForce GTX 1080 Ti50.52 tok/s11GBMeasured
NVIDIA L450.45 tok/s24GBMeasured
NVIDIA GeForce GTX 1660 Ti49.29 tok/s6GBMeasured
NVIDIA GeForce GTX 1660 Super48.64 tok/s6GBMeasured
AMD Radeon RX 670048.5 tok/s10GBEstimated
NVIDIA GeForce RTX 505046.6 tok/s8GBEstimated
Intel Arc A770 Limited Edition46.3 tok/s16GBEstimated
NVIDIA TITAN Xp (power capped)46.06 tok/s12GBMeasured
AMD Radeon RX 760044.1 tok/s8GBEstimated
NVIDIA RTX 2000 Ada Generation43.08 tok/s16GBMeasured
NVIDIA GeForce RTX 305041.4 tok/s8GBEstimated
Intel Arc A75037.0 tok/s8GBEstimated
NVIDIA GeForce GTX 108035.98 tok/s8GBMeasured
NVIDIA GeForce GTX 1070 Ti35.9 tok/s8GBEstimated
NVIDIA T435.01 tok/s16GBMeasured
NVIDIA GeForce GTX 166034.37 tok/s6GBMeasured
Intel Arc Pro A6033.6 tok/s12GBEstimated

Tap any column to sort. Measured = we rented and ran this card ourselves. Estimated = interpolated against our measured anchors, never blended silently.

Because this workload is bandwidth-bound, the ranking above tracks memory bandwidth far more closely than core counts or price. A card with fewer tensor cores and faster memory will beat a card with the opposite.

That's the reason we run twelve workloads instead of publishing one score. A GPU isn't fast or slow. It's fast at some things and gated out of others, and which of those matters depends entirely on what you're actually going to run.

Our verdict

NVIDIA B300 tops our Llama 3.1 8B leaderboard at 287.2 tok/s (measured), 755% of the way clear of the slowest card that still fits. This is a bandwidth workload: buy memory speed, not tensor cores.

FAQ

What is the fastest GPU for Llama 3.1 8B?
NVIDIA B300, at 287.2 tok/s on our bench, a first-party measurement. It carries 288GB of VRAM. Of the 102 cards we have Llama 3.1 8B data for, 100 can run it at all.
How much VRAM do I need for Llama 3.1 8B?
~8GB at Q4_K_M. Almost everything clears it, which is exactly what makes it a good yardstick.
Why does the Llama 3.1 8B ranking look different from your other benchmarks?
Because this workload is bandwidth-bound, the ranking above tracks memory bandwidth far more closely than core counts or price. A card with fewer tensor cores and faster memory will beat a card with the opposite. That's why we publish twelve separate workloads rather than one blended score, the ordering genuinely changes depending on the job.
Are these Llama 3.1 8B numbers measured or estimated?
Both, and every row says which. 71 of the 102 cards here were rented and run by us on the same harness. The remainder are anchored estimates interpolated per workload against those measurements. We never blend the two silently, if a row says Estimated, we have not run that card.
Can I rent a GPU to run Llama 3.1 8B instead of buying one?
Yes, and for the cards at the top of this leaderboard it's the only realistic option, most of them have no retail channel at all. It's also how we got these numbers: we rented the hardware by the hour rather than buying it. That's worth considering before you spend on a card to find out whether it's fast enough.
Why publish cards that can't run Llama 3.1 8B?
Nearly every card in our fleet clears this workload's VRAM floor, so there are no gates on this page. That's unusual, most of our suite excludes a meaningful chunk of the market.

How we test

Every ranking on this page comes from our own benchmark runs, not vendor claims. Cards marked Measured were rented and run by us; cards marked Estimated are interpolated per workload against those measured anchors and are labelled on every row, we never blend the two silently. LLMs run on llama.cpp (llama-bench) at Q4_K_M with -p 512 -n 128. Diffusion and video run on diffusers/ComfyUI at BF16, with SDXL at FP16. Each workload gets a warmup pass plus multiple timed runs (5 for small LLMs, 3 for large models and images, 2 for video); we publish the mean as the result and the minimum as the 1% low. Run-to-run variance is under 0.5%. Telemetry, power, temperature, utilisation, clocks, peak VRAM, is sampled at 1 Hz for the duration of every run. Where a model exceeds a card's VRAM we publish a hard won't-fit result rather than quietly dropping to a smaller quantisation. A card that can't run a model scores zero on it. Silently swapping precision to make a number appear would make every number on this site meaningless. All figures are single-GPU, single-stream, batch-size-1. That is the honest way to measure what one card does for one user, and it is deliberately not how a datacenter serves a model. Vendor and MLPerf figures use large batches across many GPUs and will be far higher. Neither is wrong, they answer different questions. Ours answers 'what will this card do for me'.