Qwen3 4B · 39 cards measured first-party · Updated July 2026

How Fast Does Qwen3 4B Run on Each GPU?

Qwen3 4B is the entry point: ~5GB, runs on essentially anything with a modern GPU in it, and fast enough that the bottleneck stops being the card and starts being how quickly you can read. It's the model to reach for on a laptop or an old 8GB card.

Benchmarked weights: Qwen/Qwen3-4B-GGUF

Fastest we measured
NVIDIA GeForce RTX 5090

NVIDIA GeForce RTX 5090

375.6 tok/s on Qwen3 4B. Measured on our bench. 32GB of VRAM, 575W board rating. AI Score 22.3/100 across our full 12-workload suite.

Pros
  • 375.6 tok/s on Qwen3 4B
  • 32GB, clears the Qwen3 4B floor
  • Rentable by the hour rather than bought
Cons
  • 575W board rating
  • Datacenter or workstation hardware, not a retail purchase

Best for: Qwen3 4B work where you want the ceiling gone rather than the cheapest entry.

Runner-up
NVIDIA H100 NVL

NVIDIA H100 NVL

364.7 tok/s on Qwen3 4B. Anchored estimate. 94GB of VRAM, 400W board rating. AI Score 67.0/100 across our full 12-workload suite.

Pros
  • 364.7 tok/s on Qwen3 4B
  • 94GB, clears the Qwen3 4B floor
  • Rentable by the hour rather than bought
Cons
  • 400W board rating
  • Datacenter or workstation hardware, not a retail purchase

Best for: Qwen3 4B work where you want the ceiling gone rather than the cheapest entry.

Third
NVIDIA RTX PRO 6000 Blackwell Workstation Edition

NVIDIA RTX PRO 6000 Blackwell Workstation Edition

362.81 tok/s on Qwen3 4B. Measured on our bench. 96GB of VRAM, 600W board rating. AI Score 53.1/100 across our full 12-workload suite.

Pros
  • 362.81 tok/s on Qwen3 4B
  • 96GB, clears the Qwen3 4B floor
  • Rentable by the hour rather than bought
Cons
  • 600W board rating
  • Datacenter or workstation hardware, not a retail purchase

Best for: Qwen3 4B work where you want the ceiling gone rather than the cheapest entry.

375.6tok/s
Fastest: NVIDIA GeForce RTX 5090
measured
61
Cards that run Qwen3 4B
of 61 we have data for
0
Cards that can't run it at all
published as hard gates, not omissions
530%
Fastest vs slowest that fits
375.6 vs 70.88 tok/s

Bandwidth-bound, but with a twist worth knowing: at 4B the model is so small that the fastest cards stop being fully occupied. On our measured B300 this workload sat at just 12-30% utilisation, the chip spends its time waiting rather than computing. Past a certain point, buying more GPU stops buying more tokens.

Qwen3 4B, the 12 fastest cards we have data for

NVIDIA GeForce RTX 5090
375.6 tok/s
NVIDIA H100 NVL
364.7 tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
362.81 tok/s
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
344.7 tok/s
NVIDIA B300
333.34 tok/s
NVIDIA GH200 Grace Hopper
325.5 tok/s
NVIDIA H200
318.84 tok/s
NVIDIA B200
317.93 tok/s
NVIDIA RTX PRO 6000 Blackwell Server Edition
315.72 tok/s
NVIDIA H800 80GB
310.3 tok/s
NVIDIA H100 80GB HBM3
310.26 tok/s
NVIDIA B100
302 tok/s

Single stream, batch size 1. 39 of the 61 cards on this page were measured first-party by us; the rest are anchored estimates against those measurements and are labelled in the table below.

Full Qwen3 4B leaderboard, every card that runs it

NVIDIA GeForce RTX 5090375.6 tok/s
NVIDIA H100 NVL364.7 tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition362.81 tok/s
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition344.7 tok/s
NVIDIA B300333.34 tok/s
NVIDIA GH200 Grace Hopper325.5 tok/s
NVIDIA H200318.84 tok/s
NVIDIA B200317.93 tok/s
NVIDIA RTX PRO 6000 Blackwell Server Edition315.72 tok/s
NVIDIA H800 80GB310.3 tok/s
NVIDIA H100 80GB HBM3310.26 tok/s
NVIDIA B100302.0 tok/s
NVIDIA RTX PRO 5000 Blackwell298.78 tok/s
NVIDIA GeForce RTX 4090260.54 tok/s
NVIDIA RTX 6000 Ada Generation241.33 tok/s
GeForce RTX 5070 Ti231.69 tok/s
NVIDIA RTX PRO 4500 Blackwell221.92 tok/s
GeForce RTX 5080215.89 tok/s
NVIDIA L40211.33 tok/s
NVIDIA L40S210.33 tok/s
NVIDIA GeForce RTX 3090206.67 tok/s
NVIDIA GeForce RTX 3080 Ti206.45 tok/s
NVIDIA GeForce RTX 4080201.76 tok/s
NVIDIA A100 80GB PCIe198.27 tok/s
NVIDIA A100 80GB SXM4197.0 tok/s
NVIDIA A800 80GB197.0 tok/s
NVIDIA RTX 5880 Ada Generation189.5 tok/s
NVIDIA RTX A5500187.5 tok/s
NVIDIA H100 PCIe185.2 tok/s
NVIDIA RTX A6000184.34 tok/s
NVIDIA GeForce RTX 3080184.08 tok/s
NVIDIA RTX PRO 4000 Blackwell179.58 tok/s
AMD Radeon Pro W7900178.8 tok/s
NVIDIA RTX A5000171.68 tok/s
NVIDIA RTX 5000 Ada Generation170.72 tok/s
NVIDIA A100 40GB PCIe159.3 tok/s
NVIDIA A100 40GB SXM4150.2 tok/s
NVIDIA GeForce RTX 4070149.68 tok/s
NVIDIA RTX A4500148.87 tok/s
NVIDIA GeForce RTX 2080 Super132.8 tok/s
NVIDIA A10G130.96 tok/s
AMD Radeon Pro W6800129.1 tok/s
AMD Radeon RX 6900 XT129.1 tok/s
NVIDIA Quadro RTX 8000127.01 tok/s
NVIDIA GeForce RTX 2060 Super126.4 tok/s
NVIDIA GeForce RTX 2070 SUPER126.4 tok/s
NVIDIA GeForce RTX 2070126.4 tok/s
NVIDIA GeForce RTX 2080 Founders Edition126.4 tok/s
NVIDIA Quadro RTX 5000126.4 tok/s
NVIDIA GeForce RTX 3070 Founders Edition126.36 tok/s
NVIDIA Quadro RTX 6000 (Turing)125.23 tok/s
NVIDIA RTX 4500 Ada Generation119.7 tok/s
NVIDIA GeForce RTX 3060 Ti118.21 tok/s
NVIDIA RTX A4000116.23 tok/s
NVIDIA RTX 4000 (Ada Generation)110.18 tok/s
NVIDIA GeForce RTX 3060102.04 tok/s
NVIDIA GeForce RTX 2060100.8 tok/s
NVIDIA GeForce RTX 4060 Ti96.29 tok/s
GeForce RTX 406086.38 tok/s
NVIDIA L484.35 tok/s
NVIDIA RTX 2000 Ada Generation70.88 tok/s
GPUResultVRAMSource
NVIDIA GeForce RTX 5090375.6 tok/s32GBMeasured
NVIDIA H100 NVL364.7 tok/s94GBEstimated
NVIDIA RTX PRO 6000 Blackwell Workstation Edition362.81 tok/s96GBMeasured
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition344.7 tok/s96GBEstimated
NVIDIA B300333.34 tok/s288GBMeasured
NVIDIA GH200 Grace Hopper325.5 tok/s141GBEstimated
NVIDIA H200318.84 tok/s141GBMeasured
NVIDIA B200317.93 tok/s192GBMeasured
NVIDIA RTX PRO 6000 Blackwell Server Edition315.72 tok/s96GBMeasured
NVIDIA H800 80GB310.3 tok/s80GBEstimated
NVIDIA H100 80GB HBM3310.26 tok/s80GBMeasured
NVIDIA B100302.0 tok/s192GBEstimated
NVIDIA RTX PRO 5000 Blackwell298.78 tok/s48GBMeasured
NVIDIA GeForce RTX 4090260.54 tok/s24GBMeasured
NVIDIA RTX 6000 Ada Generation241.33 tok/s48GBMeasured
GeForce RTX 5070 Ti231.69 tok/s16GBMeasured
NVIDIA RTX PRO 4500 Blackwell221.92 tok/s32GBMeasured
GeForce RTX 5080215.89 tok/s16GBMeasured
NVIDIA L40211.33 tok/s48GBMeasured
NVIDIA L40S210.33 tok/s48GBMeasured
NVIDIA GeForce RTX 3090206.67 tok/s24GBMeasured
NVIDIA GeForce RTX 3080 Ti206.45 tok/s12GBMeasured
NVIDIA GeForce RTX 4080201.76 tok/s16GBMeasured
NVIDIA A100 80GB PCIe198.27 tok/s80GBMeasured
NVIDIA A100 80GB SXM4197.0 tok/s80GBMeasured
NVIDIA A800 80GB197.0 tok/s80GBEstimated
NVIDIA RTX 5880 Ada Generation189.5 tok/s48GBEstimated
NVIDIA RTX A5500187.5 tok/s24GBEstimated
NVIDIA H100 PCIe185.2 tok/s80GBEstimated
NVIDIA RTX A6000184.34 tok/s48GBMeasured
NVIDIA GeForce RTX 3080184.08 tok/s10GBMeasured
NVIDIA RTX PRO 4000 Blackwell179.58 tok/s24GBMeasured
AMD Radeon Pro W7900178.8 tok/s48GBEstimated
NVIDIA RTX A5000171.68 tok/s24GBMeasured
NVIDIA RTX 5000 Ada Generation170.72 tok/s32GBMeasured
NVIDIA A100 40GB PCIe159.3 tok/s40GBEstimated
NVIDIA A100 40GB SXM4150.2 tok/s40GBEstimated
NVIDIA GeForce RTX 4070149.68 tok/s12GBMeasured
NVIDIA RTX A4500148.87 tok/s20GBMeasured
NVIDIA GeForce RTX 2080 Super132.8 tok/s8GBEstimated
NVIDIA A10G130.96 tok/s24GBMeasured
AMD Radeon Pro W6800129.1 tok/s32GBEstimated
AMD Radeon RX 6900 XT129.1 tok/s16GBEstimated
NVIDIA Quadro RTX 8000127.01 tok/s48GBMeasured
NVIDIA GeForce RTX 2060 Super126.4 tok/s8GBEstimated
NVIDIA GeForce RTX 2070 SUPER126.4 tok/s8GBEstimated
NVIDIA GeForce RTX 2070126.4 tok/s8GBEstimated
NVIDIA GeForce RTX 2080 Founders Edition126.4 tok/s8GBEstimated
NVIDIA Quadro RTX 5000126.4 tok/s16GBEstimated
NVIDIA GeForce RTX 3070 Founders Edition126.36 tok/s8GBMeasured
NVIDIA Quadro RTX 6000 (Turing)125.23 tok/s24GBMeasured
NVIDIA RTX 4500 Ada Generation119.7 tok/s24GBEstimated
NVIDIA GeForce RTX 3060 Ti118.21 tok/s8GBMeasured
NVIDIA RTX A4000116.23 tok/s16GBMeasured
NVIDIA RTX 4000 (Ada Generation)110.18 tok/s20GBMeasured
NVIDIA GeForce RTX 3060102.04 tok/s12GBMeasured
NVIDIA GeForce RTX 2060100.8 tok/s6GBEstimated
NVIDIA GeForce RTX 4060 Ti96.29 tok/s16GBMeasured
GeForce RTX 406086.38 tok/s8GBMeasured
NVIDIA L484.35 tok/s24GBMeasured
NVIDIA RTX 2000 Ada Generation70.88 tok/s16GBMeasured

Tap any column to sort. Measured = we rented and ran this card ourselves. Estimated = interpolated against our measured anchors, never blended silently.

Because this workload is bandwidth-bound, the ranking above tracks memory bandwidth far more closely than core counts or price. A card with fewer tensor cores and faster memory will beat a card with the opposite.

That's the reason we run twelve workloads instead of publishing one score. A GPU isn't fast or slow. It's fast at some things and gated out of others, and which of those matters depends entirely on what you're actually going to run.

Our verdict

NVIDIA GeForce RTX 5090 tops our Qwen3 4B leaderboard at 375.6 tok/s (measured), 530% of the way clear of the slowest card that still fits. This is a bandwidth workload: buy memory speed, not tensor cores.

FAQ

What is the fastest GPU for Qwen3 4B?
NVIDIA GeForce RTX 5090, at 375.6 tok/s on our bench, a first-party measurement. It carries 32GB of VRAM. Of the 61 cards we have Qwen3 4B data for, 61 can run it at all.
How much VRAM do I need for Qwen3 4B?
~5GB at Q4_K_M. If a card can't run this, it can't run anything in our LLM ladder.
Why does the Qwen3 4B ranking look different from your other benchmarks?
Because this workload is bandwidth-bound, the ranking above tracks memory bandwidth far more closely than core counts or price. A card with fewer tensor cores and faster memory will beat a card with the opposite. That's why we publish twelve separate workloads rather than one blended score, the ordering genuinely changes depending on the job.
Are these Qwen3 4B numbers measured or estimated?
Both, and every row says which. 39 of the 61 cards here were rented and run by us on the same harness. The remainder are anchored estimates interpolated per workload against those measurements. We never blend the two silently, if a row says Estimated, we have not run that card.
Can I rent a GPU to run Qwen3 4B instead of buying one?
Yes, and for the cards at the top of this leaderboard it's the only realistic option, most of them have no retail channel at all. It's also how we got these numbers: we rented the hardware by the hour rather than buying it. That's worth considering before you spend on a card to find out whether it's fast enough.
Why publish cards that can't run Qwen3 4B?
Nearly every card in our fleet clears this workload's VRAM floor, so there are no gates on this page. That's unusual, most of our suite excludes a meaningful chunk of the market.

How we test

Every ranking on this page comes from our own benchmark runs, not vendor claims. Cards marked Measured were rented and run by us; cards marked Estimated are interpolated per workload against those measured anchors and are labelled on every row, we never blend the two silently. LLMs run on llama.cpp (llama-bench) at Q4_K_M with -p 512 -n 128. Diffusion and video run on diffusers/ComfyUI at BF16, with SDXL at FP16. Each workload gets a warmup pass plus multiple timed runs (5 for small LLMs, 3 for large models and images, 2 for video); we publish the mean as the result and the minimum as the 1% low. Run-to-run variance is under 0.5%. Telemetry, power, temperature, utilisation, clocks, peak VRAM, is sampled at 1 Hz for the duration of every run. Where a model exceeds a card's VRAM we publish a hard won't-fit result rather than quietly dropping to a smaller quantisation. A card that can't run a model scores zero on it. Silently swapping precision to make a number appear would make every number on this site meaningless. All figures are single-GPU, single-stream, batch-size-1. That is the honest way to measure what one card does for one user, and it is deliberately not how a datacenter serves a model. Vendor and MLPerf figures use large batches across many GPUs and will be far higher. Neither is wrong, they answer different questions. Ours answers 'what will this card do for me'.