Qwen3 32B · 56 cards measured first-party · Updated October 2026

How Fast Does Qwen3 32B Run on Each GPU?

Qwen3 32B is the largest model most people can realistically run at home. It needs roughly 20GB at Q4_K_M, which clears a 24GB consumer card with room to spare and stops a 16GB card dead. That makes it the most useful single benchmark on this site for anyone choosing between the two.

Benchmarked weights: unsloth/Qwen3-32B-GGUF

Fastest we measured
NVIDIA B300

NVIDIA B300

83.68 tok/s on Qwen3 32B, the ceiling. Measured on our bench. 288GB of VRAM, $40,000 at launch.

Pros
  • 83.68 tok/s on Qwen3 32B
  • 288GB, clears the Qwen3 32B floor
  • Rentable by the hour rather than bought
Cons
  • 1400W board rating
  • Datacenter or workstation hardware, not a retail purchase
Best consumer card
NVIDIA GeForce RTX 5090

NVIDIA GeForce RTX 5090

71.15 tok/s on Qwen3 32B, fastest card you can buy at retail. Measured on our bench. 32GB of VRAM, $1,999 at launch.

Pros
  • 71.15 tok/s on Qwen3 32B
  • 32GB, clears the Qwen3 32B floor
Cons
  • 575W board rating
Cheapest card that runs it
AMD Radeon RX 7900 XTX

AMD Radeon RX 7900 XTX

41.4 tok/s on Qwen3 32B, lowest launch price that still fits. Anchored estimate. 24GB of VRAM, $999 at launch.

Pros
  • 41.4 tok/s on Qwen3 32B
  • 24GB, clears the Qwen3 32B floor
Cons
  • 355W board rating
Best value
NVIDIA GeForce RTX 4090

NVIDIA GeForce RTX 4090

44.28 tok/s on Qwen3 32B, most speed per dollar. Measured on our bench. 24GB of VRAM, $1,599 at launch. That is 27.69 tok/s per $1,000 of launch price.

Pros
  • 44.28 tok/s on Qwen3 32B
  • 24GB, clears the Qwen3 32B floor
Cons
  • 450W board rating
83.68tok/s
Fastest: NVIDIA B300
measured
42
Cards that run Qwen3 32B
of 101 we have data for
59
Cards that can't run it at all
published as hard gates, not omissions
574%
Fastest vs slowest that fits
83.68 vs 12.42 tok/s

Like every LLM in our suite, 32B token generation is bandwidth-bound. The card reads ~20GB of weights per token, so the ranking follows memory bandwidth and largely ignores compute. A card with half the tensor cores and the same bandwidth will land in roughly the same place.

Qwen3 32B: speed on every GPU we have data for

NVIDIA B300
83.68 tok/s
NVIDIA B200
78.56 tok/s
NVIDIA GH200 Grace Hopper
78.2 tok/s
NVIDIA H200
76.58 tok/s
NVIDIA B100
74.6 tok/s
NVIDIA H800 80GB
74.1 tok/s
NVIDIA H100 80GB HBM3
74.07 tok/s
NVIDIA GeForce RTX 5090
71.15 tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
70.13 tok/s
NVIDIA H100 NVL
64.06 tok/s
NVIDIA RTX PRO 6000 Blackwell Server Edition
63.9 tok/s
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
59.97 tok/s
NVIDIA H100 PCIe
54.34 tok/s
NVIDIA RTX PRO 5000 Blackwell
54.27 tok/s
NVIDIA A100 80GB SXM4
45.53 tok/s

Top 15 shown; 27 more cards in the full table below.

Single stream, batch size 1. 39 of the 61 cards on this page were measured first-party by us; the rest are anchored estimates against those measurements and are labelled in the table below.

Efficiency: tok/s per 100W drawn

NVIDIA GeForce RTX 5090
85.93 tok/s / 100W
NVIDIA H200
64.52 tok/s / 100W
NVIDIA RTX PRO 4500 Blackwell
47.91 tok/s / 100W
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
35.15 tok/s / 100W
NVIDIA H100 80GB HBM3
34.06 tok/s / 100W
NVIDIA RTX PRO 5000 Blackwell
32.71 tok/s / 100W
NVIDIA RTX PRO 6000 Blackwell Server Edition
29.33 tok/s / 100W
NVIDIA A100 80GB PCIe
28.99 tok/s / 100W
NVIDIA A100 80GB SXM4
28.76 tok/s / 100W
NVIDIA GeForce RTX 3090
27.92 tok/s / 100W
NVIDIA GeForce RTX 4090
27.02 tok/s / 100W
NVIDIA RTX A6000
25.46 tok/s / 100W
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
25.27 tok/s / 100W
NVIDIA H100 PCIe
24.51 tok/s / 100W
NVIDIA RTX PRO 4000 Blackwell
24.32 tok/s / 100W

Top 15 shown; 17 more cards in the full table below.

Power is the average pulled during the run, sampled at 1Hz. The fastest card is often not the one here, and for anything left running this is the number that shows up on the bill.

Value: tok/s per $1,000 of MSRP

NVIDIA GeForce RTX 5090
35.59 tok/s / $1k
NVIDIA GeForce RTX 4090
27.69 tok/s / $1k
NVIDIA GeForce RTX 3090
25.32 tok/s / $1k
NVIDIA GeForce RTX 3090 Ti
21.1 tok/s / $1k
NVIDIA RTX PRO 4000 Blackwell
18.53 tok/s / $1k
NVIDIA RTX PRO 4500 Blackwell
14.47 tok/s / $1k
NVIDIA RTX A5000
13.55 tok/s / $1k
NVIDIA RTX PRO 5000 Blackwell
12.06 tok/s / $1k
NVIDIA Titan RTX
11.35 tok/s / $1k
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
8.19 tok/s / $1k
NVIDIA A10G
7.87 tok/s / $1k
NVIDIA RTX PRO 6000 Blackwell Server Edition
7.46 tok/s / $1k
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
7 tok/s / $1k
NVIDIA RTX A6000
6.92 tok/s / $1k
NVIDIA RTX 5000 Ada Generation
6.47 tok/s / $1k

Top 15 shown; 17 more cards in the full table below.

Launch price, not street price, so it ages. A speed leaderboard always crowns the most expensive card; this is the counterweight.

Won't fit, Qwen3 32B gates these cards outright

AMD Radeon RX 7900 XT20GB
NVIDIA RTX 4000 (Ada Generation)20GB
NVIDIA RTX A450020GB
AMD Radeon RX 6800 XT16GB
AMD Radeon RX 680016GB
AMD Radeon RX 6900 XT16GB
AMD Radeon RX 6950 XT16GB
AMD Radeon RX 7800 XT16GB
AMD Radeon RX 9070 XT16GB
AMD Radeon RX 907016GB
NVIDIA GeForce RTX 4060 Ti 16GB16GB
NVIDIA GeForce RTX 4070 Ti Super16GB
GeForce RTX 4080 Super16GB
NVIDIA GeForce RTX 408016GB
GeForce RTX 5060 Ti16GB
GeForce RTX 5070 Ti16GB
GeForce RTX 508016GB
Intel Arc A770 Limited Edition16GB
NVIDIA Quadro RTX 500016GB
NVIDIA RTX 2000 Ada Generation16GB
NVIDIA RTX A400016GB
NVIDIA T416GB
AMD Radeon RX 7700 XT12GB
NVIDIA GeForce RTX 306012GB
NVIDIA GeForce RTX 3080 Ti12GB
NVIDIA GeForce RTX 4070 Super12GB
NVIDIA GeForce RTX 4070 Ti12GB
NVIDIA GeForce RTX 407012GB
GeForce RTX 507012GB
Intel Arc B58012GB
Intel Arc Pro A6012GB
NVIDIA TITAN V12GB
NVIDIA TITAN X (Pascal)12GB
NVIDIA TITAN Xp12GB
GeForce GTX 1080 Ti11GB
NVIDIA GeForce RTX 2080 Ti Founders Edition11GB
AMD Radeon RX 670010GB
NVIDIA GeForce RTX 308010GB
AMD Radeon RX 76008GB
NVIDIA GeForce GTX 1070 Ti8GB
NVIDIA GeForce GTX 10808GB
NVIDIA GeForce RTX 2060 Super8GB
NVIDIA GeForce RTX 2070 SUPER8GB
NVIDIA GeForce RTX 20708GB
NVIDIA GeForce RTX 2080 Super8GB
NVIDIA GeForce RTX 2080 Founders Edition8GB
NVIDIA GeForce RTX 30508GB
NVIDIA GeForce RTX 3060 Ti8GB
NVIDIA GeForce RTX 3070 Ti8GB
NVIDIA GeForce RTX 3070 Founders Edition8GB
GeForce RTX 40608GB
NVIDIA GeForce RTX 50508GB
NVIDIA GeForce RTX 50608GB
Intel Arc A7508GB
NVIDIA GeForce GTX 1660 Super6GB
NVIDIA GeForce GTX 1660 Ti6GB
NVIDIA GeForce GTX 16606GB
NVIDIA GeForce RTX 20606GB
NVIDIA RTX A20006GB
GPUVRAMWhy it fails
AMD Radeon RX 7900 XT20GBNeeds ~23GB VRAM
NVIDIA RTX 4000 (Ada Generation)20GBNeeds ~23GB VRAM
NVIDIA RTX A450020GBNeeds ~23GB VRAM
AMD Radeon RX 6800 XT16GBNeeds ~20GB VRAM
AMD Radeon RX 680016GBNeeds ~20GB VRAM
AMD Radeon RX 6900 XT16GBNeeds ~20GB VRAM
AMD Radeon RX 6950 XT16GBNeeds ~20GB VRAM
AMD Radeon RX 7800 XT16GBNeeds ~20GB VRAM
AMD Radeon RX 9070 XT16GBNeeds ~20GB VRAM
AMD Radeon RX 907016GBNeeds ~20GB VRAM
NVIDIA GeForce RTX 4060 Ti 16GB16GBNeeds ~23GB VRAM
NVIDIA GeForce RTX 4070 Ti Super16GBNeeds ~23GB VRAM
GeForce RTX 4080 Super16GBNeeds ~23GB VRAM
NVIDIA GeForce RTX 408016GBNeeds ~23GB VRAM
GeForce RTX 5060 Ti16GBNeeds ~23GB VRAM
GeForce RTX 5070 Ti16GBNeeds ~23GB VRAM
GeForce RTX 508016GBNeeds ~23GB VRAM
Intel Arc A770 Limited Edition16GBNeeds ~20GB VRAM
NVIDIA Quadro RTX 500016GBNeeds ~20GB VRAM
NVIDIA RTX 2000 Ada Generation16GBNeeds ~23GB VRAM
NVIDIA RTX A400016GBNeeds ~23GB VRAM
NVIDIA T416GBNeeds ~23GB VRAM
AMD Radeon RX 7700 XT12GBNeeds ~20GB VRAM
NVIDIA GeForce RTX 306012GBNeeds ~23GB VRAM
NVIDIA GeForce RTX 3080 Ti12GBNeeds ~23GB VRAM
NVIDIA GeForce RTX 4070 Super12GBNeeds ~23GB VRAM
NVIDIA GeForce RTX 4070 Ti12GBNeeds ~20GB VRAM
NVIDIA GeForce RTX 407012GBNeeds ~23GB VRAM
GeForce RTX 507012GBNeeds ~23GB VRAM
Intel Arc B58012GBNeeds ~20GB VRAM
Intel Arc Pro A6012GBNeeds ~20GB VRAM
NVIDIA TITAN V12GBNeeds ~20GB VRAM
NVIDIA TITAN X (Pascal)12GBNeeds ~20GB VRAM
NVIDIA TITAN Xp12GBNeeds ~20GB VRAM
GeForce GTX 1080 Ti11GBNeeds ~20GB VRAM
NVIDIA GeForce RTX 2080 Ti Founders Edition11GBNeeds ~23GB VRAM
AMD Radeon RX 670010GBNeeds ~20GB VRAM
NVIDIA GeForce RTX 308010GBNeeds ~23GB VRAM
AMD Radeon RX 76008GBNeeds ~20GB VRAM
NVIDIA GeForce GTX 1070 Ti8GBNeeds ~20GB VRAM
NVIDIA GeForce GTX 10808GBNeeds ~20GB VRAM
NVIDIA GeForce RTX 2060 Super8GBNeeds ~20GB VRAM
NVIDIA GeForce RTX 2070 SUPER8GBNeeds ~20GB VRAM
NVIDIA GeForce RTX 20708GBNeeds ~20GB VRAM
NVIDIA GeForce RTX 2080 Super8GBNeeds ~20GB VRAM
NVIDIA GeForce RTX 2080 Founders Edition8GBNeeds ~20GB VRAM
NVIDIA GeForce RTX 30508GBNeeds ~20GB VRAM
NVIDIA GeForce RTX 3060 Ti8GBNeeds ~23GB VRAM
NVIDIA GeForce RTX 3070 Ti8GBNeeds ~23GB VRAM
NVIDIA GeForce RTX 3070 Founders Edition8GBNeeds ~23GB VRAM
GeForce RTX 40608GBNeeds ~23GB VRAM
NVIDIA GeForce RTX 50508GBNeeds ~20GB VRAM
NVIDIA GeForce RTX 50608GBNeeds ~20GB VRAM
Intel Arc A7508GBNeeds ~20GB VRAM
NVIDIA GeForce GTX 1660 Super6GBNeeds ~20GB VRAM
NVIDIA GeForce GTX 1660 Ti6GBNeeds ~20GB VRAM
NVIDIA GeForce GTX 16606GBNeeds ~20GB VRAM
NVIDIA GeForce RTX 20606GBNeeds ~20GB VRAM
NVIDIA RTX A20006GBNeeds ~23GB VRAM

Showing 14 of 26. No driver update fixes a VRAM ceiling.

Full Qwen3 32B leaderboard, every card that runs it

NVIDIA B30083.68 tok/s
NVIDIA B20078.56 tok/s
NVIDIA GH200 Grace Hopper78.2 tok/s
NVIDIA H20076.58 tok/s
NVIDIA B10074.6 tok/s
NVIDIA H800 80GB74.1 tok/s
NVIDIA H100 80GB HBM374.07 tok/s
NVIDIA GeForce RTX 509071.15 tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition70.13 tok/s
NVIDIA H100 NVL64.06 tok/s
NVIDIA RTX PRO 6000 Blackwell Server Edition63.9 tok/s
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition59.97 tok/s
NVIDIA H100 PCIe54.34 tok/s
NVIDIA RTX PRO 5000 Blackwell54.27 tok/s
NVIDIA A100 80GB SXM445.53 tok/s
NVIDIA A800 80GB45.5 tok/s
NVIDIA GeForce RTX 409044.28 tok/s
NVIDIA A100 40GB SXM443.8 tok/s
NVIDIA A100 80GB PCIe43.77 tok/s
NVIDIA GeForce RTX 3090 Ti42.18 tok/s
NVIDIA A100 40GB PCIe41.5 tok/s
AMD Radeon RX 7900 XTX41.4 tok/s
NVIDIA RTX 5880 Ada Generation40.79 tok/s
NVIDIA RTX 6000 Ada Generation39.87 tok/s
NVIDIA GeForce RTX 309037.95 tok/s
NVIDIA RTX PRO 4500 Blackwell37.61 tok/s
NVIDIA L40S34.44 tok/s
NVIDIA L4034.08 tok/s
NVIDIA RTX A600032.18 tok/s
AMD Radeon Pro W790032.0 tok/s
NVIDIA RTX A550030.5 tok/s
NVIDIA RTX A500030.49 tok/s
NVIDIA Titan RTX28.36 tok/s
NVIDIA RTX PRO 4000 Blackwell27.8 tok/s
NVIDIA A4026.89 tok/s
NVIDIA RTX 5000 Ada Generation25.88 tok/s
NVIDIA A10G22.04 tok/s
NVIDIA Quadro RTX 800021.98 tok/s
AMD Radeon Pro W780020.2 tok/s
NVIDIA RTX 4500 Ada Generation18.2 tok/s
AMD Radeon Pro W680018.1 tok/s
NVIDIA L412.42 tok/s
GPUResultVRAMSource
NVIDIA B30083.68 tok/s288GBMeasured
NVIDIA B20078.56 tok/s192GBMeasured
NVIDIA GH200 Grace Hopper78.2 tok/s141GBEstimated
NVIDIA H20076.58 tok/s141GBMeasured
NVIDIA B10074.6 tok/s192GBEstimated
NVIDIA H800 80GB74.1 tok/s80GBEstimated
NVIDIA H100 80GB HBM374.07 tok/s80GBMeasured
NVIDIA GeForce RTX 509071.15 tok/s32GBMeasured
NVIDIA RTX PRO 6000 Blackwell Workstation Edition70.13 tok/s96GBMeasured
NVIDIA H100 NVL64.06 tok/s94GBMeasured
NVIDIA RTX PRO 6000 Blackwell Server Edition63.9 tok/s96GBMeasured
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition59.97 tok/s96GBMeasured
NVIDIA H100 PCIe54.34 tok/s80GBMeasured
NVIDIA RTX PRO 5000 Blackwell54.27 tok/s48GBMeasured
NVIDIA A100 80GB SXM445.53 tok/s80GBMeasured
NVIDIA A800 80GB45.5 tok/s80GBEstimated
NVIDIA GeForce RTX 409044.28 tok/s24GBMeasured
NVIDIA A100 40GB SXM443.8 tok/s40GBMeasured
NVIDIA A100 80GB PCIe43.77 tok/s80GBMeasured
NVIDIA GeForce RTX 3090 Ti42.18 tok/s24GBMeasured
NVIDIA A100 40GB PCIe41.5 tok/s40GBMeasured
AMD Radeon RX 7900 XTX41.4 tok/s24GBEstimated
NVIDIA RTX 5880 Ada Generation40.79 tok/s48GBMeasured
NVIDIA RTX 6000 Ada Generation39.87 tok/s48GBMeasured
NVIDIA GeForce RTX 309037.95 tok/s24GBMeasured
NVIDIA RTX PRO 4500 Blackwell37.61 tok/s32GBMeasured
NVIDIA L40S34.44 tok/s48GBMeasured
NVIDIA L4034.08 tok/s48GBMeasured
NVIDIA RTX A600032.18 tok/s48GBMeasured
AMD Radeon Pro W790032.0 tok/s48GBEstimated
NVIDIA RTX A550030.5 tok/s24GBEstimated
NVIDIA RTX A500030.49 tok/s24GBMeasured
NVIDIA Titan RTX28.36 tok/s24GBMeasured
NVIDIA RTX PRO 4000 Blackwell27.8 tok/s24GBMeasured
NVIDIA A4026.89 tok/s48GBMeasured
NVIDIA RTX 5000 Ada Generation25.88 tok/s32GBMeasured
NVIDIA A10G22.04 tok/s24GBMeasured
NVIDIA Quadro RTX 800021.98 tok/s48GBMeasured
AMD Radeon Pro W780020.2 tok/s32GBEstimated
NVIDIA RTX 4500 Ada Generation18.2 tok/s24GBEstimated
AMD Radeon Pro W680018.1 tok/s32GBEstimated
NVIDIA L412.42 tok/s24GBMeasured

Tap any column to sort. Measured = we rented and ran this card ourselves. Estimated = interpolated against our measured anchors, never blended silently.

Because this workload is bandwidth-bound, the ranking above tracks memory bandwidth far more closely than core counts or price. A card with fewer tensor cores and faster memory will beat a card with the opposite.

That's the reason we run twelve workloads instead of publishing one score. A GPU isn't fast or slow. It's fast at some things and gated out of others, and which of those matters depends entirely on what you're actually going to run.

Our verdict

NVIDIA B300 tops our Qwen3 32B leaderboard at 83.68 tok/s (measured), 574% of the way clear of the slowest card that still fits. But the number that decides most purchases isn't on the chart. It's the 59 cards that can't run Qwen3 32B at all. This is a bandwidth workload: buy memory speed, not tensor cores.

FAQ

What is the fastest GPU for Qwen3 32B?
NVIDIA B300, at 83.68 tok/s on our bench, a first-party measurement. It carries 288GB of VRAM. Of the 101 cards we have Qwen3 32B data for, 42 can run it at all.
How much VRAM do I need for Qwen3 32B?
~20GB at Q4_K_M. This is the cliff that decides the 16GB-versus-24GB question, and it's why we treat VRAM as the first spec and speed as the second.
Why does the Qwen3 32B ranking look different from your other benchmarks?
Because this workload is bandwidth-bound, the ranking above tracks memory bandwidth far more closely than core counts or price. A card with fewer tensor cores and faster memory will beat a card with the opposite. That's why we publish twelve separate workloads rather than one blended score, the ordering genuinely changes depending on the job.
Are these Qwen3 32B numbers measured or estimated?
Both, and every row says which. 56 of the 101 cards here were rented and run by us on the same harness. The remainder are anchored estimates interpolated per workload against those measurements. We never blend the two silently, if a row says Estimated, we have not run that card.
Can I rent a GPU to run Qwen3 32B instead of buying one?
Yes, and for the cards at the top of this leaderboard it's the only realistic option, most of them have no retail channel at all. It's also how we got these numbers: we rented the hardware by the hour rather than buying it. That's worth considering before you spend on a card to find out whether it's fast enough.
Why publish cards that can't run Qwen3 32B?
Because it's the most useful thing we know. A card that can't load a model doesn't run it slowly, it doesn't run it. Most benchmark sites leave that as a blank cell or quietly drop to a smaller quantisation to produce a number. We publish it as a hard gate and score it zero, because 'this card cannot do the thing you want' is the answer to the question you were actually asking.

How we test

Every ranking on this page comes from our own benchmark runs, not vendor claims. Cards marked Measured were rented and run by us; cards marked Estimated are interpolated per workload against those measured anchors and are labelled on every row, we never blend the two silently. LLMs run on llama.cpp (llama-bench) at Q4_K_M with -p 512 -n 128. Diffusion and video run on diffusers/ComfyUI at BF16, with SDXL at FP16. Each workload gets a warmup pass plus multiple timed runs (5 for small LLMs, 3 for large models and images, 2 for video); we publish the mean as the result and the minimum as the 1% low. Run-to-run variance is under 0.5%. Telemetry, power, temperature, utilisation, clocks, peak VRAM, is sampled at 1 Hz for the duration of every run. Where a model exceeds a card's VRAM we publish a hard won't-fit result rather than quietly dropping to a smaller quantisation. A card that can't run a model scores zero on it. Silently swapping precision to make a number appear would make every number on this site meaningless. All figures are single-GPU, single-stream, batch-size-1. That is the honest way to measure what one card does for one user, and it is deliberately not how a datacenter serves a model. Vendor and MLPerf figures use large batches across many GPUs and will be far higher. Neither is wrong, they answer different questions. Ours answers 'what will this card do for me'.