Qwen3-Coder 30B A3B · 13 GPUs measured first-party · llama.cpp Q4_K_M · Updated October 2026
Qwen3 Coder 30B-A3B is the coding-tuned version of the 30B MoE, and the fastest 'big' model in our entire database: 318 tok/s on the RTX PRO 6000 Blackwell at just 110W. Same mixture-of-experts economics as its sibling (30B in memory, ~3B active per token), same ~20GB measured floor, measured on 13 GPUs with our pinned llama.cpp harness.
Benchmarked weights: bartowski/Qwen_Qwen3-Coder-30B-A3B-Instruct-GGUF

365.8 tok/s on Qwen3-Coder 30B A3B, the ceiling. Measured on our bench. 32GB of VRAM, $1,999 at launch.

271.0 tok/s on Qwen3-Coder 30B A3B, fastest card you can buy at retail. Measured on our bench. 24GB of VRAM, $1,599 at launch.

219.1 tok/s on Qwen3-Coder 30B A3B, lowest launch price that still fits. Measured on our bench. 24GB of VRAM, $1,499 at launch.

317.7 tok/s on Qwen3-Coder 30B A3B, most speed per dollar. Measured on our bench. 96GB of VRAM, $8,565 at launch. That is 37.09 tok/s per $1,000 of launch price.
What GPU Do You Need for Qwen3-Coder 30B A3B?, tok/s by GPU
Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.
Efficiency: tok/s per 100W drawn
Power is the average pulled during the run, sampled at 1Hz. The fastest card is often not the one here, and for anything left running this is the number that shows up on the bill.
Value: tok/s per $1,000 of MSRP
Launch price, not street price, so it ages. A speed leaderboard always crowns the most expensive card; this is the counterweight.
Qwen3-Coder 30B A3B. Measured generation speed by GPU
| GPU | tok/s | Prompt t/s | tok/W | Avg power |
|---|---|---|---|---|
| NVIDIA GeForce RTX 5090 | 365.8 | 11452.4 | 1.86 | 197.1 W |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 317.7 | 9786.1 | 2.87 | 110.5 W |
| NVIDIA H200 | 297.9 | 7287.6 | 2.39 | 124.4 W |
| NVIDIA H100 80GB HBM3 | 296.2 | 7235.9 | 2.33 | 127.4 W |
| NVIDIA B300 | 284.9 | 4674.1 | 1.11 | 256.2 W |
| NVIDIA B200 | 277.7 | 7658.5 | 0.98 | 282.0 W |
| NVIDIA GeForce RTX 4090 | 271 | 9315.9 | 1.85 | 146.4 W |
| NVIDIA GeForce RTX 3090 | 219.1 | 4756.4 | 0.95 | 229.7 W |
| NVIDIA L40S | 219 | 8054.6 | 1.65 | 132.6 W |
| NVIDIA A100 80GB SXM4 | 182.3 | 3754.5 | 1.67 | 109.0 W |
| NVIDIA A100 40GB SXM4 | 173.5 | 3666.4 | 1.88 | 92.1 W |
| NVIDIA A10G | 142.8 | 2689.2 | 1.47 | 96.9 W |
| NVIDIA L4 | 99.14 | 2424.5 | 1.87 | 53.0 W |
The local coding agent, solved. Coding is the workload where generation speed matters most: agent loops burn thousands of tokens on edits, diffs and retries, and a slow model makes the whole loop unusable. That's why this model is significant: 30B-class code quality at 300+ tok/s means a local coding agent that keeps up with you. If I were building a local-first coding setup on a 24GB card today, this is the model I'd build it around.
The efficiency story is the headline. 318 tok/s at 110.5W is 2.87 tokens per watt, the best big-model efficiency figure we've measured. The H200 and H100 sit just behind (298 and 296 tok/s), and even the A10G holds 145 tok/s, which still beats every dense 32B result in our data on any GPU. The ~20GB floor puts it on 24GB consumer cards; pair it with the smaller Qwen2.5-Coder 7B as a fast autocomplete model and you have a full local coding stack on one card.
About Qwen3-Coder 30B A3B. Qwen3-Coder 30B A3B: from Qwen, 31B parameters, on Hugging Face since July 2025, Apache 2.0 licence. 9,073,846 downloads in the last 30 days and 4 community quantizations.
How it compares. H100 80GB HBM3: Qwen3-Coder 30B A3B 296.2 tok/s, Qwen3 30B A3B 283.8 (31B), Qwen3 30B A3B Instruct 2507 299.8 (31B), GLM-4.7-Flash 186.1 (31B), Gemma 4 31B 76.59 (31B). 1 of 4 beat Qwen3-Coder 30B A3B here.
Cost on a rented GPU. 1M generated tokens of Qwen3-Coder 30B A3B: $0.15 on a RTX 3090 ($0.12/hr, 76 min), $0.30 on a RTX 5090 ($0.39/hr, 46 min, 1.9x the cost).
Qwen3-Coder 30B A3B: cost per 1M generated tokens on rented GPUs
| GPU | Cheapest rate | Speed (tok/s) | Cost per 1M generated tokens |
|---|---|---|---|
| NVIDIA GeForce RTX 3090 | $0.12/hr | 219.1 | $0.15 |
| NVIDIA GeForce RTX 5090 | $0.39/hr | 365.8 | $0.30 |
| NVIDIA GeForce RTX 4090 | $0.34/hr | 271 | $0.34 |
| NVIDIA A100 40GB SXM4 | $0.47/hr | 173.5 | $0.76 |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | $1.08/hr | 317.7 | $0.94 |
| NVIDIA L40S | $0.79/hr | 219 | $1.00 |
| NVIDIA L4 | $0.44/hr | 99.14 | $1.23 |
| NVIDIA A100 80GB SXM4 | $0.95/hr | 182.3 | $1.44 |
| NVIDIA H100 80GB HBM3 | $2.14/hr | 296.2 | $2.00 |
| NVIDIA H200 | $3.59/hr | 297.9 | $3.35 |
| NVIDIA B200 | $5.98/hr | 277.7 | $5.98 |
| NVIDIA B300 | $6.94/hr | 284.9 | $6.77 |
Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.
Speed tiers for Qwen3-Coder 30B A3B. 30+ tok/s: 13 (RTX 5090, RTX 4090, RTX 3090). 30 tok/s is roughly where replies outpace reading.
Reading your prompt. Before Qwen3-Coder 30B A3B writes anything it reads the input: 11452.4 tok/s on the RTX 5090 (0.3s for a 4,000-token prompt), 9315.9 on the RTX 4090 (0.4s), 2424.5 on the L4 (1.6s). Long documents and big code files feel this number more than the generation speed.
VRAM for Qwen3-Coder 30B A3B. Measured peak 17.8GB, so 24GB is the smallest common card size; smallest card it ran on: RTX 4090 (24GB). With long context: Q4_K_M 20GB (tested), Q2_K 13GB, Q3_K_M 16GB, Q5_K_M 24GB, Q6_K 28GB.
Power on Qwen3-Coder 30B A3B. Most efficient: A100 40GB SXM4, 92W, 0.15 kWh per 1M generated tokens. Hungriest: B200, 282W, 0.28 kWh. At $0.15/kWh: $0.022 per 1M generated tokens.
Qwen3 Coder 30B-A3B: 318 tok/s at 110W, the fastest big model we've benchmarked, full stop. MoE speed makes local coding agents actually viable, and the ~20GB floor fits any 24GB card. Our measurements, our harness, logged power on every run.