Llama 3.3 70B · 39 cards measured first-party · Updated July 2026

How Fast Does Llama 3.3 70B Run on Each GPU?

Llama 3.3 70B is the model that separates a serious AI machine from an expensive one. At Q4_K_M it needs roughly 42GB of VRAM, and that single number disqualifies more of the GPU market than any other figure in our suite, including every consumer card ever made. We ran it on every GPU that could hold it and published a hard gate on every GPU that couldn't.

Benchmarked weights: bartowski/Llama-3.3-70B-Instruct-GGUF

Fastest we measured
NVIDIA H100 NVL

NVIDIA H100 NVL

48.2 tok/s on Llama 3.3 70B. Anchored estimate. 94GB of VRAM, 400W board rating. AI Score 67.0/100 across our full 12-workload suite.

Pros
  • 48.2 tok/s on Llama 3.3 70B
  • 94GB, clears the Llama 3.3 70B floor
  • Rentable by the hour rather than bought
Cons
  • 400W board rating
  • Datacenter or workstation hardware, not a retail purchase

Best for: Llama 3.3 70B work where you want the ceiling gone rather than the cheapest entry.

Runner-up
NVIDIA B300

NVIDIA B300

47.97 tok/s on Llama 3.3 70B. Measured on our bench. 288GB of VRAM, 1400W board rating. AI Score 93.8/100 across our full 12-workload suite.

Pros
  • 47.97 tok/s on Llama 3.3 70B
  • 288GB, clears the Llama 3.3 70B floor
  • Rentable by the hour rather than bought
Cons
  • 1400W board rating
  • Datacenter or workstation hardware, not a retail purchase

Best for: Llama 3.3 70B work where you want the ceiling gone rather than the cheapest entry.

Third
NVIDIA B200

NVIDIA B200

44.54 tok/s on Llama 3.3 70B. Measured on our bench. 192GB of VRAM, 1000W board rating. AI Score 78.0/100 across our full 12-workload suite.

Pros
  • 44.54 tok/s on Llama 3.3 70B
  • 192GB, clears the Llama 3.3 70B floor
  • Rentable by the hour rather than bought
Cons
  • 1000W board rating
  • Datacenter or workstation hardware, not a retail purchase

Best for: Llama 3.3 70B work where you want the ceiling gone rather than the cheapest entry.

48.2tok/s
Fastest: NVIDIA H100 NVL
anchored estimate
20
Cards that run Llama 3.3 70B
of 61 we have data for
41
Cards that can't run it at all
published as hard gates, not omissions
357%
Fastest vs slowest that fits
48.2 vs 13.5 tok/s

Token generation on a 70B is bound by memory bandwidth, not compute. The model has to be read out of VRAM once per token, so the ceiling is how fast the card can move 42GB of weights, not how many tensor cores it has. That's why the ranking below tracks bandwidth almost perfectly and ignores core counts, and why an HBM card with modest compute buries a consumer flagship that can't even load the thing.

Llama 3.3 70B, the 12 fastest cards we have data for

NVIDIA H100 NVL
48.2 tok/s
NVIDIA B300
47.97 tok/s
NVIDIA B200
44.54 tok/s
NVIDIA GH200 Grace Hopper
43.5 tok/s
NVIDIA H200
42.66 tok/s
NVIDIA B100
42.3 tok/s
NVIDIA H100 80GB HBM3
41 tok/s
NVIDIA H800 80GB
41 tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
34.87 tok/s
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
33.1 tok/s
NVIDIA RTX PRO 6000 Blackwell Server Edition
32.08 tok/s
NVIDIA RTX PRO 5000 Blackwell
26.29 tok/s

Single stream, batch size 1. 39 of the 61 cards on this page were measured first-party by us; the rest are anchored estimates against those measurements and are labelled in the table below.

Won't fit, Llama 3.3 70B gates these cards outright

NVIDIA L4048GB
NVIDIA L40S48GB
NVIDIA Quadro RTX 800048GB
NVIDIA A100 40GB PCIe40GB
NVIDIA A100 40GB SXM440GB
AMD Radeon Pro W680032GB
NVIDIA GeForce RTX 509032GB
NVIDIA RTX 5000 Ada Generation32GB
NVIDIA RTX PRO 4500 Blackwell32GB
NVIDIA GeForce RTX 309024GB
NVIDIA GeForce RTX 409024GB
NVIDIA A10G24GB
NVIDIA L424GB
NVIDIA Quadro RTX 6000 (Turing)24GB
GPUVRAMWhy it fails
NVIDIA L4048GBrequires ~46GB VRAM
NVIDIA L40S48GBrequires ~46GB VRAM
NVIDIA Quadro RTX 800048GBrequires ~46GB VRAM
NVIDIA A100 40GB PCIe40GBNeeds needs ~42GB VRAM
NVIDIA A100 40GB SXM440GBNeeds needs ~42GB VRAM
AMD Radeon Pro W680032GBNeeds needs ~42GB VRAM
NVIDIA GeForce RTX 509032GBrequires ~46GB VRAM
NVIDIA RTX 5000 Ada Generation32GBrequires ~46GB VRAM
NVIDIA RTX PRO 4500 Blackwell32GBrequires ~46GB VRAM
NVIDIA GeForce RTX 309024GBrequires ~46GB VRAM
NVIDIA GeForce RTX 409024GBrequires ~46GB VRAM
NVIDIA A10G24GBrequires ~46GB VRAM
NVIDIA L424GBrequires ~46GB VRAM
NVIDIA Quadro RTX 6000 (Turing)24GBrequires ~46GB VRAM

Showing 14 of 41. No driver update fixes a VRAM ceiling.

Full Llama 3.3 70B leaderboard, every card that runs it

NVIDIA H100 NVL48.2 tok/s
NVIDIA B30047.97 tok/s
NVIDIA B20044.54 tok/s
NVIDIA GH200 Grace Hopper43.5 tok/s
NVIDIA H20042.66 tok/s
NVIDIA B10042.3 tok/s
NVIDIA H100 80GB HBM341.0 tok/s
NVIDIA H800 80GB41.0 tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition34.87 tok/s
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition33.1 tok/s
NVIDIA RTX PRO 6000 Blackwell Server Edition32.08 tok/s
NVIDIA RTX PRO 5000 Blackwell26.29 tok/s
NVIDIA H100 PCIe24.5 tok/s
NVIDIA A100 80GB SXM424.4 tok/s
NVIDIA A800 80GB24.4 tok/s
NVIDIA A100 80GB PCIe22.89 tok/s
NVIDIA RTX 6000 Ada Generation18.4 tok/s
NVIDIA RTX A600015.73 tok/s
AMD Radeon Pro W790014.1 tok/s
NVIDIA RTX 5880 Ada Generation13.5 tok/s
GPUResultVRAMSource
NVIDIA H100 NVL48.2 tok/s94GBEstimated
NVIDIA B30047.97 tok/s288GBMeasured
NVIDIA B20044.54 tok/s192GBMeasured
NVIDIA GH200 Grace Hopper43.5 tok/s141GBEstimated
NVIDIA H20042.66 tok/s141GBMeasured
NVIDIA B10042.3 tok/s192GBEstimated
NVIDIA H100 80GB HBM341.0 tok/s80GBMeasured
NVIDIA H800 80GB41.0 tok/s80GBEstimated
NVIDIA RTX PRO 6000 Blackwell Workstation Edition34.87 tok/s96GBMeasured
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition33.1 tok/s96GBEstimated
NVIDIA RTX PRO 6000 Blackwell Server Edition32.08 tok/s96GBMeasured
NVIDIA RTX PRO 5000 Blackwell26.29 tok/s48GBMeasured
NVIDIA H100 PCIe24.5 tok/s80GBEstimated
NVIDIA A100 80GB SXM424.4 tok/s80GBMeasured
NVIDIA A800 80GB24.4 tok/s80GBEstimated
NVIDIA A100 80GB PCIe22.89 tok/s80GBMeasured
NVIDIA RTX 6000 Ada Generation18.4 tok/s48GBMeasured
NVIDIA RTX A600015.73 tok/s48GBMeasured
AMD Radeon Pro W790014.1 tok/s48GBEstimated
NVIDIA RTX 5880 Ada Generation13.5 tok/s48GBEstimated

Tap any column to sort. Measured = we rented and ran this card ourselves. Estimated = interpolated against our measured anchors, never blended silently.

Because this workload is bandwidth-bound, the ranking above tracks memory bandwidth far more closely than core counts or price. A card with fewer tensor cores and faster memory will beat a card with the opposite.

That's the reason we run twelve workloads instead of publishing one score. A GPU isn't fast or slow. It's fast at some things and gated out of others, and which of those matters depends entirely on what you're actually going to run.

Our verdict

NVIDIA H100 NVL tops our Llama 3.3 70B leaderboard at 48.2 tok/s (anchored estimate), 357% of the way clear of the slowest card that still fits. But the number that decides most purchases isn't on the chart. It's the 41 cards that can't run Llama 3.3 70B at all. This is a bandwidth workload: buy memory speed, not tensor cores.

FAQ

What is the fastest GPU for Llama 3.3 70B?
NVIDIA H100 NVL, at 48.2 tok/s on our bench, an anchored estimate against our measured cards. It carries 94GB of VRAM. Of the 61 cards we have Llama 3.3 70B data for, 20 can run it at all.
How much VRAM do I need for Llama 3.3 70B?
42GB at Q4_K_M is the floor. There is no driver update, no optimisation and no setting that gets a 32GB card past it, the weights either fit or they don't. Cards below the line score zero on this workload in our AI Score, because a card that can't run the job doesn't get partial credit for being quick at the jobs it can.
Why does the Llama 3.3 70B ranking look different from your other benchmarks?
Because this workload is bandwidth-bound, the ranking above tracks memory bandwidth far more closely than core counts or price. A card with fewer tensor cores and faster memory will beat a card with the opposite. That's why we publish twelve separate workloads rather than one blended score, the ordering genuinely changes depending on the job.
Are these Llama 3.3 70B numbers measured or estimated?
Both, and every row says which. 39 of the 61 cards here were rented and run by us on the same harness. The remainder are anchored estimates interpolated per workload against those measurements. We never blend the two silently, if a row says Estimated, we have not run that card.
Can I rent a GPU to run Llama 3.3 70B instead of buying one?
Yes, and for the cards at the top of this leaderboard it's the only realistic option, most of them have no retail channel at all. It's also how we got these numbers: we rented the hardware by the hour rather than buying it. That's worth considering before you spend on a card to find out whether it's fast enough.
Why publish cards that can't run Llama 3.3 70B?
Because it's the most useful thing we know. A card that can't load a model doesn't run it slowly, it doesn't run it. Most benchmark sites leave that as a blank cell or quietly drop to a smaller quantisation to produce a number. We publish it as a hard gate and score it zero, because 'this card cannot do the thing you want' is the answer to the question you were actually asking.

How we test

Every ranking on this page comes from our own benchmark runs, not vendor claims. Cards marked Measured were rented and run by us; cards marked Estimated are interpolated per workload against those measured anchors and are labelled on every row, we never blend the two silently. LLMs run on llama.cpp (llama-bench) at Q4_K_M with -p 512 -n 128. Diffusion and video run on diffusers/ComfyUI at BF16, with SDXL at FP16. Each workload gets a warmup pass plus multiple timed runs (5 for small LLMs, 3 for large models and images, 2 for video); we publish the mean as the result and the minimum as the 1% low. Run-to-run variance is under 0.5%. Telemetry, power, temperature, utilisation, clocks, peak VRAM, is sampled at 1 Hz for the duration of every run. Where a model exceeds a card's VRAM we publish a hard won't-fit result rather than quietly dropping to a smaller quantisation. A card that can't run a model scores zero on it. Silently swapping precision to make a number appear would make every number on this site meaningless. All figures are single-GPU, single-stream, batch-size-1. That is the honest way to measure what one card does for one user, and it is deliberately not how a datacenter serves a model. Vendor and MLPerf figures use large batches across many GPUs and will be far higher. Neither is wrong, they answer different questions. Ours answers 'what will this card do for me'.