Qwen3 4B · 71 cards measured first-party · Updated October 2026

How Fast Does Qwen3 4B Run on Each GPU?

Qwen3 4B is the entry point: ~5GB, runs on essentially anything with a modern GPU in it, and fast enough that the bottleneck stops being the card and starts being how quickly you can read. It's the model to reach for on a laptop or an old 8GB card.

Benchmarked weights: Qwen/Qwen3-4B-GGUF

Fastest we measured
NVIDIA GeForce RTX 5090

NVIDIA GeForce RTX 5090

375.6 tok/s on Qwen3 4B, the ceiling. Measured on our bench. 32GB of VRAM, $1,999 at launch.

Pros
  • 375.6 tok/s on Qwen3 4B
  • 32GB, clears the Qwen3 4B floor
Cons
  • 575W board rating

Best for: Qwen3 4B work where you want the ceiling gone rather than the cheapest entry.

Best consumer card
NVIDIA GeForce RTX 4090

NVIDIA GeForce RTX 4090

260.5 tok/s on Qwen3 4B, fastest card you can buy at retail. Measured on our bench. 24GB of VRAM, $1,599 at launch.

Pros
  • 260.5 tok/s on Qwen3 4B
  • 24GB, clears the Qwen3 4B floor
Cons
  • 450W board rating
Cheapest card that runs it
Intel Arc B580

Intel Arc B580

94.9 tok/s on Qwen3 4B, lowest launch price that still fits. Anchored estimate. 12GB of VRAM, $179 at launch.

Pros
  • 94.9 tok/s on Qwen3 4B
  • 12GB, clears the Qwen3 4B floor
Cons
  • 190W board rating
Best value
NVIDIA GeForce RTX 5060

NVIDIA GeForce RTX 5060

119.5 tok/s on Qwen3 4B, most speed per dollar. Measured on our bench. 8GB of VRAM, $249 at launch. That is 480.1 tok/s per $1,000 of launch price.

Pros
  • 119.5 tok/s on Qwen3 4B
  • 8GB, clears the Qwen3 4B floor
Cons
  • 145W board rating
375.6tok/s
Fastest: NVIDIA GeForce RTX 5090
measured
102
Cards that run Qwen3 4B
of 102 we have data for
0
Cards that can't run it at all
published as hard gates, not omissions
585%
Fastest vs slowest that fits
375.6 vs 54.8 tok/s

Bandwidth-bound, but with a twist worth knowing: at 4B the model is so small that the fastest cards stop being fully occupied. On our measured B300 this workload sat at just 12-30% utilisation, the chip spends its time waiting rather than computing. Past a certain point, buying more GPU stops buying more tokens.

Qwen3 4B: speed on every GPU we have data for

NVIDIA GeForce RTX 5090
375.6 tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
362.8 tok/s
NVIDIA B300
333.3 tok/s
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
330.9 tok/s
NVIDIA GH200 Grace Hopper
325.5 tok/s
NVIDIA H200
318.8 tok/s
NVIDIA B200
317.9 tok/s
NVIDIA RTX PRO 6000 Blackwell Server Edition
315.7 tok/s
NVIDIA H800 80GB
310.3 tok/s
NVIDIA H100 80GB HBM3
310.3 tok/s
NVIDIA B100
302 tok/s
NVIDIA RTX PRO 5000 Blackwell
298.8 tok/s
NVIDIA H100 NVL
282.9 tok/s
NVIDIA GeForce RTX 4090
260.5 tok/s
NVIDIA H100 PCIe
245.9 tok/s

Top 15 shown; 87 more cards in the full table below.

Single stream, batch size 1. 39 of the 61 cards on this page were measured first-party by us; the rest are anchored estimates against those measurements and are labelled in the table below.

Efficiency: tok/s per 100W drawn

NVIDIA GeForce RTX 5090
431.23 tok/s / 100W
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
389.7 tok/s / 100W
NVIDIA H200
297.15 tok/s / 100W
NVIDIA RTX PRO 4500 Blackwell
287.09 tok/s / 100W
GeForce RTX 5080
251.33 tok/s / 100W
NVIDIA RTX PRO 5000 Blackwell
228.6 tok/s / 100W
NVIDIA GeForce RTX 4090
224.22 tok/s / 100W
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
222.66 tok/s / 100W
NVIDIA RTX 2000 Ada Generation
202.51 tok/s / 100W
NVIDIA RTX PRO 4000 Blackwell
187.45 tok/s / 100W
NVIDIA H100 80GB HBM3
183.91 tok/s / 100W
NVIDIA RTX PRO 6000 Blackwell Server Edition
183.13 tok/s / 100W
NVIDIA GeForce RTX 4080
177.29 tok/s / 100W
NVIDIA A100 80GB PCIe
173.62 tok/s / 100W
GeForce RTX 4080 Super
172.06 tok/s / 100W

Top 15 shown; 55 more cards in the full table below.

Power is the average pulled during the run, sampled at 1Hz. The fastest card is often not the one here, and for anything left running this is the number that shows up on the bill.

Value: tok/s per $1,000 of MSRP

NVIDIA GeForce RTX 5060
480.08 tok/s / $1k
NVIDIA GeForce GTX 1660 Super
348.73 tok/s / $1k
GeForce RTX 5070
328.25 tok/s / $1k
GeForce RTX 5060 Ti
314.59 tok/s / $1k
NVIDIA GeForce RTX 3060
310.15 tok/s / $1k
GeForce RTX 5070 Ti
309.33 tok/s / $1k
NVIDIA GeForce RTX 3060 Ti
296.27 tok/s / $1k
GeForce RTX 4060
288.9 tok/s / $1k
NVIDIA GeForce GTX 1660 Ti
288.71 tok/s / $1k
NVIDIA GeForce GTX 1660
267.17 tok/s / $1k
NVIDIA GeForce RTX 3080
263.35 tok/s / $1k
NVIDIA GeForce RTX 3070 Ti
258.18 tok/s / $1k
NVIDIA GeForce RTX 4070 Super
255.16 tok/s / $1k
NVIDIA GeForce RTX 3070 Founders Edition
253.23 tok/s / $1k
NVIDIA GeForce RTX 4070
249.88 tok/s / $1k

Top 15 shown; 56 more cards in the full table below.

Launch price, not street price, so it ages. A speed leaderboard always crowns the most expensive card; this is the counterweight.

Full Qwen3 4B leaderboard, every card that runs it

NVIDIA GeForce RTX 5090375.6 tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition362.8 tok/s
NVIDIA B300333.3 tok/s
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition330.9 tok/s
NVIDIA GH200 Grace Hopper325.5 tok/s
NVIDIA H200318.8 tok/s
NVIDIA B200317.9 tok/s
NVIDIA RTX PRO 6000 Blackwell Server Edition315.7 tok/s
NVIDIA H800 80GB310.3 tok/s
NVIDIA H100 80GB HBM3310.3 tok/s
NVIDIA B100302.0 tok/s
NVIDIA RTX PRO 5000 Blackwell298.8 tok/s
NVIDIA H100 NVL282.9 tok/s
NVIDIA GeForce RTX 4090260.5 tok/s
NVIDIA H100 PCIe245.9 tok/s
NVIDIA RTX 5880 Ada Generation242.1 tok/s
NVIDIA RTX 6000 Ada Generation241.3 tok/s
GeForce RTX 5070 Ti231.7 tok/s
NVIDIA GeForce RTX 3090 Ti227.5 tok/s
NVIDIA RTX PRO 4500 Blackwell221.9 tok/s
GeForce RTX 5080215.9 tok/s
AMD Radeon RX 7900 XTX213.9 tok/s
NVIDIA L40S212.1 tok/s
NVIDIA L40211.3 tok/s
NVIDIA GeForce RTX 3090206.7 tok/s
NVIDIA GeForce RTX 3080 Ti206.4 tok/s
NVIDIA GeForce RTX 4080201.8 tok/s
NVIDIA A100 80GB PCIe198.3 tok/s
GeForce RTX 4080 Super197.5 tok/s
NVIDIA A100 40GB PCIe197.3 tok/s
NVIDIA A100 80GB SXM4197.0 tok/s
NVIDIA A800 80GB197.0 tok/s
NVIDIA A100 40GB SXM4193.1 tok/s
NVIDIA GeForce RTX 4070 Ti Super187.8 tok/s
NVIDIA RTX A5500187.5 tok/s
NVIDIA RTX A6000184.3 tok/s
NVIDIA GeForce RTX 3080184.1 tok/s
AMD Radeon RX 7900 XT183.3 tok/s
GeForce RTX 5070180.2 tok/s
NVIDIA RTX PRO 4000 Blackwell179.6 tok/s
AMD Radeon Pro W7900178.8 tok/s
NVIDIA RTX A5000171.7 tok/s
NVIDIA RTX 5000 Ada Generation170.7 tok/s
NVIDIA A40160.8 tok/s
AMD Radeon RX 9070 XT158.5 tok/s
NVIDIA Titan RTX158.1 tok/s
AMD Radeon RX 7800 XT156.2 tok/s
NVIDIA GeForce RTX 3070 Ti154.7 tok/s
NVIDIA GeForce RTX 4070 Super152.8 tok/s
AMD Radeon RX 9070152.2 tok/s
NVIDIA GeForce RTX 4070 Ti151.7 tok/s
NVIDIA GeForce RTX 4070149.7 tok/s
NVIDIA RTX A4500148.9 tok/s
NVIDIA GeForce RTX 2080 Ti Founders Edition148.5 tok/s
NVIDIA TITAN V141.1 tok/s
AMD Radeon RX 6900 XT135.6 tok/s
GeForce RTX 5060 Ti135.0 tok/s
NVIDIA GeForce RTX 2080 Super132.8 tok/s
NVIDIA RTX 4500 Ada Generation132.2 tok/s
AMD Radeon Pro W7800130.0 tok/s
AMD Radeon RX 6950 XT130.0 tok/s
NVIDIA A10G129.9 tok/s
AMD Radeon Pro W6800129.1 tok/s
AMD Radeon RX 6800 XT129.1 tok/s
AMD Radeon RX 6800129.1 tok/s
NVIDIA Quadro RTX 8000127.0 tok/s
NVIDIA GeForce RTX 2080 Founders Edition126.4 tok/s
NVIDIA GeForce RTX 3070 Founders Edition126.4 tok/s
NVIDIA Quadro RTX 6000 (Turing)125.2 tok/s
NVIDIA GeForce RTX 2070 SUPER119.7 tok/s
NVIDIA GeForce RTX 5060119.5 tok/s
NVIDIA Quadro RTX 5000118.3 tok/s
NVIDIA GeForce RTX 3060 Ti118.2 tok/s
NVIDIA RTX A4000116.2 tok/s
NVIDIA RTX 4000 (Ada Generation)110.2 tok/s
NVIDIA GeForce RTX 2070 (power capped)107.3 tok/s
NVIDIA GeForce RTX 3060102.0 tok/s
AMD Radeon RX 7700 XT101.7 tok/s
NVIDIA GeForce RTX 2060100.8 tok/s
NVIDIA TITAN X (Pascal)100.1 tok/s
NVIDIA GeForce RTX 2060 Super (power capped)98.72 tok/s
NVIDIA GeForce RTX 4060 Ti 16GB96.29 tok/s
Intel Arc B58094.9 tok/s
GeForce RTX 406086.38 tok/s
NVIDIA L484.0 tok/s
NVIDIA GeForce GTX 1660 Ti80.55 tok/s
NVIDIA GeForce GTX 1660 Super79.86 tok/s
AMD Radeon RX 670078.5 tok/s
GeForce GTX 1080 Ti78.25 tok/s
NVIDIA TITAN Xp (power capped)75.21 tok/s
NVIDIA GeForce RTX 505073.2 tok/s
Intel Arc A770 Limited Edition71.1 tok/s
NVIDIA RTX 2000 Ada Generation70.88 tok/s
AMD Radeon RX 760070.4 tok/s
NVIDIA RTX A200068.47 tok/s
NVIDIA T467.71 tok/s
NVIDIA GeForce RTX 305063.7 tok/s
Intel Arc A75059.5 tok/s
NVIDIA GeForce GTX 108059.12 tok/s
NVIDIA GeForce GTX 166058.51 tok/s
NVIDIA GeForce GTX 1070 Ti56.4 tok/s
Intel Arc Pro A6054.8 tok/s
GPUResultVRAMSource
NVIDIA GeForce RTX 5090375.6 tok/s32GBMeasured
NVIDIA RTX PRO 6000 Blackwell Workstation Edition362.8 tok/s96GBMeasured
NVIDIA B300333.3 tok/s288GBMeasured
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition330.9 tok/s96GBMeasured
NVIDIA GH200 Grace Hopper325.5 tok/s141GBEstimated
NVIDIA H200318.8 tok/s141GBMeasured
NVIDIA B200317.9 tok/s192GBMeasured
NVIDIA RTX PRO 6000 Blackwell Server Edition315.7 tok/s96GBMeasured
NVIDIA H800 80GB310.3 tok/s80GBEstimated
NVIDIA H100 80GB HBM3310.3 tok/s80GBMeasured
NVIDIA B100302.0 tok/s192GBEstimated
NVIDIA RTX PRO 5000 Blackwell298.8 tok/s48GBMeasured
NVIDIA H100 NVL282.9 tok/s94GBMeasured
NVIDIA GeForce RTX 4090260.5 tok/s24GBMeasured
NVIDIA H100 PCIe245.9 tok/s80GBMeasured
NVIDIA RTX 5880 Ada Generation242.1 tok/s48GBMeasured
NVIDIA RTX 6000 Ada Generation241.3 tok/s48GBMeasured
GeForce RTX 5070 Ti231.7 tok/s16GBMeasured
NVIDIA GeForce RTX 3090 Ti227.5 tok/s24GBMeasured
NVIDIA RTX PRO 4500 Blackwell221.9 tok/s32GBMeasured
GeForce RTX 5080215.9 tok/s16GBMeasured
AMD Radeon RX 7900 XTX213.9 tok/s24GBEstimated
NVIDIA L40S212.1 tok/s48GBMeasured
NVIDIA L40211.3 tok/s48GBMeasured
NVIDIA GeForce RTX 3090206.7 tok/s24GBMeasured
NVIDIA GeForce RTX 3080 Ti206.4 tok/s12GBMeasured
NVIDIA GeForce RTX 4080201.8 tok/s16GBMeasured
NVIDIA A100 80GB PCIe198.3 tok/s80GBMeasured
GeForce RTX 4080 Super197.5 tok/s16GBMeasured
NVIDIA A100 40GB PCIe197.3 tok/s40GBMeasured
NVIDIA A100 80GB SXM4197.0 tok/s80GBMeasured
NVIDIA A800 80GB197.0 tok/s80GBEstimated
NVIDIA A100 40GB SXM4193.1 tok/s40GBMeasured
NVIDIA GeForce RTX 4070 Ti Super187.8 tok/s16GBMeasured
NVIDIA RTX A5500187.5 tok/s24GBEstimated
NVIDIA RTX A6000184.3 tok/s48GBMeasured
NVIDIA GeForce RTX 3080184.1 tok/s10GBMeasured
AMD Radeon RX 7900 XT183.3 tok/s20GBEstimated
GeForce RTX 5070180.2 tok/s12GBMeasured
NVIDIA RTX PRO 4000 Blackwell179.6 tok/s24GBMeasured
AMD Radeon Pro W7900178.8 tok/s48GBEstimated
NVIDIA RTX A5000171.7 tok/s24GBMeasured
NVIDIA RTX 5000 Ada Generation170.7 tok/s32GBMeasured
NVIDIA A40160.8 tok/s48GBMeasured
AMD Radeon RX 9070 XT158.5 tok/s16GBEstimated
NVIDIA Titan RTX158.1 tok/s24GBMeasured
AMD Radeon RX 7800 XT156.2 tok/s16GBEstimated
NVIDIA GeForce RTX 3070 Ti154.7 tok/s8GBMeasured
NVIDIA GeForce RTX 4070 Super152.8 tok/s12GBMeasured
AMD Radeon RX 9070152.2 tok/s16GBEstimated
NVIDIA GeForce RTX 4070 Ti151.7 tok/s12GBMeasured
NVIDIA GeForce RTX 4070149.7 tok/s12GBMeasured
NVIDIA RTX A4500148.9 tok/s20GBMeasured
NVIDIA GeForce RTX 2080 Ti Founders Edition148.5 tok/s11GBMeasured
NVIDIA TITAN V141.1 tok/s12GBMeasured
AMD Radeon RX 6900 XT135.6 tok/s16GBEstimated
GeForce RTX 5060 Ti135.0 tok/s16GBMeasured
NVIDIA GeForce RTX 2080 Super132.8 tok/s8GBEstimated
NVIDIA RTX 4500 Ada Generation132.2 tok/s24GBMeasured
AMD Radeon Pro W7800130.0 tok/s32GBEstimated
AMD Radeon RX 6950 XT130.0 tok/s16GBEstimated
NVIDIA A10G129.9 tok/s24GBMeasured
AMD Radeon Pro W6800129.1 tok/s32GBEstimated
AMD Radeon RX 6800 XT129.1 tok/s16GBEstimated
AMD Radeon RX 6800129.1 tok/s16GBEstimated
NVIDIA Quadro RTX 8000127.0 tok/s48GBMeasured
NVIDIA GeForce RTX 2080 Founders Edition126.4 tok/s8GBEstimated
NVIDIA GeForce RTX 3070 Founders Edition126.4 tok/s8GBMeasured
NVIDIA Quadro RTX 6000 (Turing)125.2 tok/s24GBMeasured
NVIDIA GeForce RTX 2070 SUPER119.7 tok/s8GBMeasured
NVIDIA GeForce RTX 5060119.5 tok/s8GBMeasured
NVIDIA Quadro RTX 5000118.3 tok/s16GBMeasured
NVIDIA GeForce RTX 3060 Ti118.2 tok/s8GBMeasured
NVIDIA RTX A4000116.2 tok/s16GBMeasured
NVIDIA RTX 4000 (Ada Generation)110.2 tok/s20GBMeasured
NVIDIA GeForce RTX 2070 (power capped)107.3 tok/s8GBMeasured
NVIDIA GeForce RTX 3060102.0 tok/s12GBMeasured
AMD Radeon RX 7700 XT101.7 tok/s12GBEstimated
NVIDIA GeForce RTX 2060100.8 tok/s6GBEstimated
NVIDIA TITAN X (Pascal)100.1 tok/s12GBEstimated
NVIDIA GeForce RTX 2060 Super (power capped)98.72 tok/s8GBMeasured
NVIDIA GeForce RTX 4060 Ti 16GB96.29 tok/s16GBMeasured
Intel Arc B58094.9 tok/s12GBEstimated
GeForce RTX 406086.38 tok/s8GBMeasured
NVIDIA L484.0 tok/s24GBMeasured
NVIDIA GeForce GTX 1660 Ti80.55 tok/s6GBMeasured
NVIDIA GeForce GTX 1660 Super79.86 tok/s6GBMeasured
AMD Radeon RX 670078.5 tok/s10GBEstimated
GeForce GTX 1080 Ti78.25 tok/s11GBMeasured
NVIDIA TITAN Xp (power capped)75.21 tok/s12GBMeasured
NVIDIA GeForce RTX 505073.2 tok/s8GBEstimated
Intel Arc A770 Limited Edition71.1 tok/s16GBEstimated
NVIDIA RTX 2000 Ada Generation70.88 tok/s16GBMeasured
AMD Radeon RX 760070.4 tok/s8GBEstimated
NVIDIA RTX A200068.47 tok/s6GBMeasured
NVIDIA T467.71 tok/s16GBMeasured
NVIDIA GeForce RTX 305063.7 tok/s8GBEstimated
Intel Arc A75059.5 tok/s8GBEstimated
NVIDIA GeForce GTX 108059.12 tok/s8GBMeasured
NVIDIA GeForce GTX 166058.51 tok/s6GBMeasured
NVIDIA GeForce GTX 1070 Ti56.4 tok/s8GBEstimated
Intel Arc Pro A6054.8 tok/s12GBEstimated

Tap any column to sort. Measured = we rented and ran this card ourselves. Estimated = interpolated against our measured anchors, never blended silently.

Because this workload is bandwidth-bound, the ranking above tracks memory bandwidth far more closely than core counts or price. A card with fewer tensor cores and faster memory will beat a card with the opposite.

That's the reason we run twelve workloads instead of publishing one score. A GPU isn't fast or slow. It's fast at some things and gated out of others, and which of those matters depends entirely on what you're actually going to run.

Our verdict

NVIDIA GeForce RTX 5090 tops our Qwen3 4B leaderboard at 375.6 tok/s (measured), 585% of the way clear of the slowest card that still fits. This is a bandwidth workload: buy memory speed, not tensor cores.

FAQ

What is the fastest GPU for Qwen3 4B?
NVIDIA GeForce RTX 5090, at 375.6 tok/s on our bench, a first-party measurement. It carries 32GB of VRAM. Of the 102 cards we have Qwen3 4B data for, 102 can run it at all.
How much VRAM do I need for Qwen3 4B?
~5GB at Q4_K_M. If a card can't run this, it can't run anything in our LLM ladder.
Why does the Qwen3 4B ranking look different from your other benchmarks?
Because this workload is bandwidth-bound, the ranking above tracks memory bandwidth far more closely than core counts or price. A card with fewer tensor cores and faster memory will beat a card with the opposite. That's why we publish twelve separate workloads rather than one blended score, the ordering genuinely changes depending on the job.
Are these Qwen3 4B numbers measured or estimated?
Both, and every row says which. 71 of the 102 cards here were rented and run by us on the same harness. The remainder are anchored estimates interpolated per workload against those measurements. We never blend the two silently, if a row says Estimated, we have not run that card.
Can I rent a GPU to run Qwen3 4B instead of buying one?
Yes, and for the cards at the top of this leaderboard it's the only realistic option, most of them have no retail channel at all. It's also how we got these numbers: we rented the hardware by the hour rather than buying it. That's worth considering before you spend on a card to find out whether it's fast enough.
Why publish cards that can't run Qwen3 4B?
Nearly every card in our fleet clears this workload's VRAM floor, so there are no gates on this page. That's unusual, most of our suite excludes a meaningful chunk of the market.

How we test

Every ranking on this page comes from our own benchmark runs, not vendor claims. Cards marked Measured were rented and run by us; cards marked Estimated are interpolated per workload against those measured anchors and are labelled on every row, we never blend the two silently. LLMs run on llama.cpp (llama-bench) at Q4_K_M with -p 512 -n 128. Diffusion and video run on diffusers/ComfyUI at BF16, with SDXL at FP16. Each workload gets a warmup pass plus multiple timed runs (5 for small LLMs, 3 for large models and images, 2 for video); we publish the mean as the result and the minimum as the 1% low. Run-to-run variance is under 0.5%. Telemetry, power, temperature, utilisation, clocks, peak VRAM, is sampled at 1 Hz for the duration of every run. Where a model exceeds a card's VRAM we publish a hard won't-fit result rather than quietly dropping to a smaller quantisation. A card that can't run a model scores zero on it. Silently swapping precision to make a number appear would make every number on this site meaningless. All figures are single-GPU, single-stream, batch-size-1. That is the honest way to measure what one card does for one user, and it is deliberately not how a datacenter serves a model. Vendor and MLPerf figures use large batches across many GPUs and will be far higher. Neither is wrong, they answer different questions. Ours answers 'what will this card do for me'.