Qwen3 4B · 71 cards measured first-party · Updated October 2026
Qwen3 4B is the entry point: ~5GB, runs on essentially anything with a modern GPU in it, and fast enough that the bottleneck stops being the card and starts being how quickly you can read. It's the model to reach for on a laptop or an old 8GB card.
Benchmarked weights: Qwen/Qwen3-4B-GGUF

375.6 tok/s on Qwen3 4B, the ceiling. Measured on our bench. 32GB of VRAM, $1,999 at launch.
Best for: Qwen3 4B work where you want the ceiling gone rather than the cheapest entry.

260.5 tok/s on Qwen3 4B, fastest card you can buy at retail. Measured on our bench. 24GB of VRAM, $1,599 at launch.

94.9 tok/s on Qwen3 4B, lowest launch price that still fits. Anchored estimate. 12GB of VRAM, $179 at launch.

119.5 tok/s on Qwen3 4B, most speed per dollar. Measured on our bench. 8GB of VRAM, $249 at launch. That is 480.1 tok/s per $1,000 of launch price.
Bandwidth-bound, but with a twist worth knowing: at 4B the model is so small that the fastest cards stop being fully occupied. On our measured B300 this workload sat at just 12-30% utilisation, the chip spends its time waiting rather than computing. Past a certain point, buying more GPU stops buying more tokens.
Qwen3 4B: speed on every GPU we have data for
Top 15 shown; 87 more cards in the full table below.
Single stream, batch size 1. 39 of the 61 cards on this page were measured first-party by us; the rest are anchored estimates against those measurements and are labelled in the table below.
Efficiency: tok/s per 100W drawn
Top 15 shown; 55 more cards in the full table below.
Power is the average pulled during the run, sampled at 1Hz. The fastest card is often not the one here, and for anything left running this is the number that shows up on the bill.
Value: tok/s per $1,000 of MSRP
Top 15 shown; 56 more cards in the full table below.
Launch price, not street price, so it ages. A speed leaderboard always crowns the most expensive card; this is the counterweight.
Full Qwen3 4B leaderboard, every card that runs it
| GPU | Result | VRAM | Source |
|---|---|---|---|
| NVIDIA GeForce RTX 5090 | 375.6 tok/s | 32GB | Measured |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 362.8 tok/s | 96GB | Measured |
| NVIDIA B300 | 333.3 tok/s | 288GB | Measured |
| NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition | 330.9 tok/s | 96GB | Measured |
| NVIDIA GH200 Grace Hopper | 325.5 tok/s | 141GB | Estimated |
| NVIDIA H200 | 318.8 tok/s | 141GB | Measured |
| NVIDIA B200 | 317.9 tok/s | 192GB | Measured |
| NVIDIA RTX PRO 6000 Blackwell Server Edition | 315.7 tok/s | 96GB | Measured |
| NVIDIA H800 80GB | 310.3 tok/s | 80GB | Estimated |
| NVIDIA H100 80GB HBM3 | 310.3 tok/s | 80GB | Measured |
| NVIDIA B100 | 302.0 tok/s | 192GB | Estimated |
| NVIDIA RTX PRO 5000 Blackwell | 298.8 tok/s | 48GB | Measured |
| NVIDIA H100 NVL | 282.9 tok/s | 94GB | Measured |
| NVIDIA GeForce RTX 4090 | 260.5 tok/s | 24GB | Measured |
| NVIDIA H100 PCIe | 245.9 tok/s | 80GB | Measured |
| NVIDIA RTX 5880 Ada Generation | 242.1 tok/s | 48GB | Measured |
| NVIDIA RTX 6000 Ada Generation | 241.3 tok/s | 48GB | Measured |
| GeForce RTX 5070 Ti | 231.7 tok/s | 16GB | Measured |
| NVIDIA GeForce RTX 3090 Ti | 227.5 tok/s | 24GB | Measured |
| NVIDIA RTX PRO 4500 Blackwell | 221.9 tok/s | 32GB | Measured |
| GeForce RTX 5080 | 215.9 tok/s | 16GB | Measured |
| AMD Radeon RX 7900 XTX | 213.9 tok/s | 24GB | Estimated |
| NVIDIA L40S | 212.1 tok/s | 48GB | Measured |
| NVIDIA L40 | 211.3 tok/s | 48GB | Measured |
| NVIDIA GeForce RTX 3090 | 206.7 tok/s | 24GB | Measured |
| NVIDIA GeForce RTX 3080 Ti | 206.4 tok/s | 12GB | Measured |
| NVIDIA GeForce RTX 4080 | 201.8 tok/s | 16GB | Measured |
| NVIDIA A100 80GB PCIe | 198.3 tok/s | 80GB | Measured |
| GeForce RTX 4080 Super | 197.5 tok/s | 16GB | Measured |
| NVIDIA A100 40GB PCIe | 197.3 tok/s | 40GB | Measured |
| NVIDIA A100 80GB SXM4 | 197.0 tok/s | 80GB | Measured |
| NVIDIA A800 80GB | 197.0 tok/s | 80GB | Estimated |
| NVIDIA A100 40GB SXM4 | 193.1 tok/s | 40GB | Measured |
| NVIDIA GeForce RTX 4070 Ti Super | 187.8 tok/s | 16GB | Measured |
| NVIDIA RTX A5500 | 187.5 tok/s | 24GB | Estimated |
| NVIDIA RTX A6000 | 184.3 tok/s | 48GB | Measured |
| NVIDIA GeForce RTX 3080 | 184.1 tok/s | 10GB | Measured |
| AMD Radeon RX 7900 XT | 183.3 tok/s | 20GB | Estimated |
| GeForce RTX 5070 | 180.2 tok/s | 12GB | Measured |
| NVIDIA RTX PRO 4000 Blackwell | 179.6 tok/s | 24GB | Measured |
| AMD Radeon Pro W7900 | 178.8 tok/s | 48GB | Estimated |
| NVIDIA RTX A5000 | 171.7 tok/s | 24GB | Measured |
| NVIDIA RTX 5000 Ada Generation | 170.7 tok/s | 32GB | Measured |
| NVIDIA A40 | 160.8 tok/s | 48GB | Measured |
| AMD Radeon RX 9070 XT | 158.5 tok/s | 16GB | Estimated |
| NVIDIA Titan RTX | 158.1 tok/s | 24GB | Measured |
| AMD Radeon RX 7800 XT | 156.2 tok/s | 16GB | Estimated |
| NVIDIA GeForce RTX 3070 Ti | 154.7 tok/s | 8GB | Measured |
| NVIDIA GeForce RTX 4070 Super | 152.8 tok/s | 12GB | Measured |
| AMD Radeon RX 9070 | 152.2 tok/s | 16GB | Estimated |
| NVIDIA GeForce RTX 4070 Ti | 151.7 tok/s | 12GB | Measured |
| NVIDIA GeForce RTX 4070 | 149.7 tok/s | 12GB | Measured |
| NVIDIA RTX A4500 | 148.9 tok/s | 20GB | Measured |
| NVIDIA GeForce RTX 2080 Ti Founders Edition | 148.5 tok/s | 11GB | Measured |
| NVIDIA TITAN V | 141.1 tok/s | 12GB | Measured |
| AMD Radeon RX 6900 XT | 135.6 tok/s | 16GB | Estimated |
| GeForce RTX 5060 Ti | 135.0 tok/s | 16GB | Measured |
| NVIDIA GeForce RTX 2080 Super | 132.8 tok/s | 8GB | Estimated |
| NVIDIA RTX 4500 Ada Generation | 132.2 tok/s | 24GB | Measured |
| AMD Radeon Pro W7800 | 130.0 tok/s | 32GB | Estimated |
| AMD Radeon RX 6950 XT | 130.0 tok/s | 16GB | Estimated |
| NVIDIA A10G | 129.9 tok/s | 24GB | Measured |
| AMD Radeon Pro W6800 | 129.1 tok/s | 32GB | Estimated |
| AMD Radeon RX 6800 XT | 129.1 tok/s | 16GB | Estimated |
| AMD Radeon RX 6800 | 129.1 tok/s | 16GB | Estimated |
| NVIDIA Quadro RTX 8000 | 127.0 tok/s | 48GB | Measured |
| NVIDIA GeForce RTX 2080 Founders Edition | 126.4 tok/s | 8GB | Estimated |
| NVIDIA GeForce RTX 3070 Founders Edition | 126.4 tok/s | 8GB | Measured |
| NVIDIA Quadro RTX 6000 (Turing) | 125.2 tok/s | 24GB | Measured |
| NVIDIA GeForce RTX 2070 SUPER | 119.7 tok/s | 8GB | Measured |
| NVIDIA GeForce RTX 5060 | 119.5 tok/s | 8GB | Measured |
| NVIDIA Quadro RTX 5000 | 118.3 tok/s | 16GB | Measured |
| NVIDIA GeForce RTX 3060 Ti | 118.2 tok/s | 8GB | Measured |
| NVIDIA RTX A4000 | 116.2 tok/s | 16GB | Measured |
| NVIDIA RTX 4000 (Ada Generation) | 110.2 tok/s | 20GB | Measured |
| NVIDIA GeForce RTX 2070 (power capped) | 107.3 tok/s | 8GB | Measured |
| NVIDIA GeForce RTX 3060 | 102.0 tok/s | 12GB | Measured |
| AMD Radeon RX 7700 XT | 101.7 tok/s | 12GB | Estimated |
| NVIDIA GeForce RTX 2060 | 100.8 tok/s | 6GB | Estimated |
| NVIDIA TITAN X (Pascal) | 100.1 tok/s | 12GB | Estimated |
| NVIDIA GeForce RTX 2060 Super (power capped) | 98.72 tok/s | 8GB | Measured |
| NVIDIA GeForce RTX 4060 Ti 16GB | 96.29 tok/s | 16GB | Measured |
| Intel Arc B580 | 94.9 tok/s | 12GB | Estimated |
| GeForce RTX 4060 | 86.38 tok/s | 8GB | Measured |
| NVIDIA L4 | 84.0 tok/s | 24GB | Measured |
| NVIDIA GeForce GTX 1660 Ti | 80.55 tok/s | 6GB | Measured |
| NVIDIA GeForce GTX 1660 Super | 79.86 tok/s | 6GB | Measured |
| AMD Radeon RX 6700 | 78.5 tok/s | 10GB | Estimated |
| GeForce GTX 1080 Ti | 78.25 tok/s | 11GB | Measured |
| NVIDIA TITAN Xp (power capped) | 75.21 tok/s | 12GB | Measured |
| NVIDIA GeForce RTX 5050 | 73.2 tok/s | 8GB | Estimated |
| Intel Arc A770 Limited Edition | 71.1 tok/s | 16GB | Estimated |
| NVIDIA RTX 2000 Ada Generation | 70.88 tok/s | 16GB | Measured |
| AMD Radeon RX 7600 | 70.4 tok/s | 8GB | Estimated |
| NVIDIA RTX A2000 | 68.47 tok/s | 6GB | Measured |
| NVIDIA T4 | 67.71 tok/s | 16GB | Measured |
| NVIDIA GeForce RTX 3050 | 63.7 tok/s | 8GB | Estimated |
| Intel Arc A750 | 59.5 tok/s | 8GB | Estimated |
| NVIDIA GeForce GTX 1080 | 59.12 tok/s | 8GB | Measured |
| NVIDIA GeForce GTX 1660 | 58.51 tok/s | 6GB | Measured |
| NVIDIA GeForce GTX 1070 Ti | 56.4 tok/s | 8GB | Estimated |
| Intel Arc Pro A60 | 54.8 tok/s | 12GB | Estimated |
Tap any column to sort. Measured = we rented and ran this card ourselves. Estimated = interpolated against our measured anchors, never blended silently.
Because this workload is bandwidth-bound, the ranking above tracks memory bandwidth far more closely than core counts or price. A card with fewer tensor cores and faster memory will beat a card with the opposite.
That's the reason we run twelve workloads instead of publishing one score. A GPU isn't fast or slow. It's fast at some things and gated out of others, and which of those matters depends entirely on what you're actually going to run.
NVIDIA GeForce RTX 5090 tops our Qwen3 4B leaderboard at 375.6 tok/s (measured), 585% of the way clear of the slowest card that still fits. This is a bandwidth workload: buy memory speed, not tensor cores.
Every ranking on this page comes from our own benchmark runs, not vendor claims. Cards marked Measured were rented and run by us; cards marked Estimated are interpolated per workload against those measured anchors and are labelled on every row, we never blend the two silently. LLMs run on llama.cpp (llama-bench) at Q4_K_M with -p 512 -n 128. Diffusion and video run on diffusers/ComfyUI at BF16, with SDXL at FP16. Each workload gets a warmup pass plus multiple timed runs (5 for small LLMs, 3 for large models and images, 2 for video); we publish the mean as the result and the minimum as the 1% low. Run-to-run variance is under 0.5%. Telemetry, power, temperature, utilisation, clocks, peak VRAM, is sampled at 1 Hz for the duration of every run. Where a model exceeds a card's VRAM we publish a hard won't-fit result rather than quietly dropping to a smaller quantisation. A card that can't run a model scores zero on it. Silently swapping precision to make a number appear would make every number on this site meaningless. All figures are single-GPU, single-stream, batch-size-1. That is the honest way to measure what one card does for one user, and it is deliberately not how a datacenter serves a model. Vendor and MLPerf figures use large batches across many GPUs and will be far higher. Neither is wrong, they answer different questions. Ours answers 'what will this card do for me'.