Qwen3-Coder 30B A3B · 13 GPUs measured first-party · llama.cpp Q4_K_M · Updated October 2026

What GPU Do You Need for Qwen3-Coder 30B A3B?

Qwen3 Coder 30B-A3B is the coding-tuned version of the 30B MoE, and the fastest 'big' model in our entire database: 318 tok/s on the RTX PRO 6000 Blackwell at just 110W. Same mixture-of-experts economics as its sibling (30B in memory, ~3B active per token), same ~20GB measured floor, measured on 13 GPUs with our pinned llama.cpp harness.

Benchmarked weights: bartowski/Qwen_Qwen3-Coder-30B-A3B-Instruct-GGUF

Fastest we measured
NVIDIA GeForce RTX 5090

NVIDIA GeForce RTX 5090

365.8 tok/s on Qwen3-Coder 30B A3B, the ceiling. Measured on our bench. 32GB of VRAM, $1,999 at launch.

Pros
  • 365.8 tok/s on Qwen3-Coder 30B A3B
  • 32GB, clears the Qwen3-Coder 30B A3B floor
Cons
  • 575W board rating
Best consumer card
NVIDIA GeForce RTX 4090

NVIDIA GeForce RTX 4090

271.0 tok/s on Qwen3-Coder 30B A3B, fastest card you can buy at retail. Measured on our bench. 24GB of VRAM, $1,599 at launch.

Pros
  • 271.0 tok/s on Qwen3-Coder 30B A3B
  • 24GB, clears the Qwen3-Coder 30B A3B floor
Cons
  • 450W board rating
Cheapest card that runs it
NVIDIA GeForce RTX 3090

NVIDIA GeForce RTX 3090

219.1 tok/s on Qwen3-Coder 30B A3B, lowest launch price that still fits. Measured on our bench. 24GB of VRAM, $1,499 at launch.

Pros
  • 219.1 tok/s on Qwen3-Coder 30B A3B
  • 24GB, clears the Qwen3-Coder 30B A3B floor
Cons
  • 350W board rating
Best value
NVIDIA RTX PRO 6000 Blackwell Workstation Edition

NVIDIA RTX PRO 6000 Blackwell Workstation Edition

317.7 tok/s on Qwen3-Coder 30B A3B, most speed per dollar. Measured on our bench. 96GB of VRAM, $8,565 at launch. That is 37.09 tok/s per $1,000 of launch price.

Pros
  • 317.7 tok/s on Qwen3-Coder 30B A3B
  • 96GB, clears the Qwen3-Coder 30B A3B floor
  • Rentable by the hour rather than bought
Cons
  • 600W board rating
  • Datacenter or workstation hardware, not a retail purchase
365.8tok/s
Fastest: NVIDIA GeForce RTX 5090
measured
10
Cards that run Qwen3-Coder 30B A3B
of 11 we have data for
1
Cards that can't run it at all
published as hard gates, not omissions
220%
Fastest vs slowest that fits
317.7 vs 99.14 tok/s

What GPU Do You Need for Qwen3-Coder 30B A3B?, tok/s by GPU

NVIDIA GeForce RTX 5090
365.8 tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
317.7 tok/s
NVIDIA H200
297.9 tok/s
NVIDIA H100 80GB HBM3
296.2 tok/s
NVIDIA B300
284.9 tok/s
NVIDIA B200
277.7 tok/s
NVIDIA GeForce RTX 4090
271 tok/s
NVIDIA GeForce RTX 3090
219.1 tok/s
NVIDIA L40S
219 tok/s
NVIDIA A100 80GB SXM4
182.3 tok/s
NVIDIA A100 40GB SXM4
173.5 tok/s
NVIDIA A10G
142.8 tok/s
NVIDIA L4
99.14 tok/s

Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.

Efficiency: tok/s per 100W drawn

NVIDIA RTX PRO 6000 Blackwell Workstation Edition
287.49 tok/s / 100W
NVIDIA H200
239.47 tok/s / 100W
NVIDIA H100 80GB HBM3
232.51 tok/s / 100W
NVIDIA A100 40GB SXM4
188.35 tok/s / 100W
NVIDIA L4
187.06 tok/s / 100W
NVIDIA GeForce RTX 5090
185.61 tok/s / 100W
NVIDIA GeForce RTX 4090
185.14 tok/s / 100W
NVIDIA A100 80GB SXM4
167.21 tok/s / 100W
NVIDIA L40S
165.19 tok/s / 100W
NVIDIA A10G
147.34 tok/s / 100W
NVIDIA B300
111.19 tok/s / 100W
NVIDIA B200
98.47 tok/s / 100W
NVIDIA GeForce RTX 3090
95.37 tok/s / 100W

Power is the average pulled during the run, sampled at 1Hz. The fastest card is often not the one here, and for anything left running this is the number that shows up on the bill.

Value: tok/s per $1,000 of MSRP

NVIDIA GeForce RTX 5090
183.01 tok/s / $1k
NVIDIA GeForce RTX 4090
169.51 tok/s / $1k
NVIDIA GeForce RTX 3090
146.14 tok/s / $1k
NVIDIA A10G
50.99 tok/s / $1k
NVIDIA L4
39.66 tok/s / $1k
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
37.09 tok/s / $1k
NVIDIA L40S
29.21 tok/s / $1k
NVIDIA A100 40GB SXM4
14.46 tok/s / $1k
NVIDIA A100 80GB SXM4
10.72 tok/s / $1k
NVIDIA H100 80GB HBM3
9.87 tok/s / $1k
NVIDIA H200
9.61 tok/s / $1k
NVIDIA B300
7.12 tok/s / $1k
NVIDIA B200
6.94 tok/s / $1k

Launch price, not street price, so it ages. A speed leaderboard always crowns the most expensive card; this is the counterweight.

Qwen3-Coder 30B A3B. Measured generation speed by GPU

NVIDIA GeForce RTX 5090365.8
NVIDIA RTX PRO 6000 Blackwell Workstation Edition317.7
NVIDIA H200297.9
NVIDIA H100 80GB HBM3296.2
NVIDIA B300284.9
NVIDIA B200277.7
NVIDIA GeForce RTX 4090271
NVIDIA GeForce RTX 3090219.1
NVIDIA L40S219
NVIDIA A100 80GB SXM4182.3
NVIDIA A100 40GB SXM4173.5
NVIDIA A10G142.8
NVIDIA L499.14
GPUtok/sPrompt t/stok/WAvg power
NVIDIA GeForce RTX 5090365.811452.41.86197.1 W
NVIDIA RTX PRO 6000 Blackwell Workstation Edition317.79786.12.87110.5 W
NVIDIA H200297.97287.62.39124.4 W
NVIDIA H100 80GB HBM3296.27235.92.33127.4 W
NVIDIA B300284.94674.11.11256.2 W
NVIDIA B200277.77658.50.98282.0 W
NVIDIA GeForce RTX 40902719315.91.85146.4 W
NVIDIA GeForce RTX 3090219.14756.40.95229.7 W
NVIDIA L40S2198054.61.65132.6 W
NVIDIA A100 80GB SXM4182.33754.51.67109.0 W
NVIDIA A100 40GB SXM4173.53666.41.8892.1 W
NVIDIA A10G142.82689.21.4796.9 W
NVIDIA L499.142424.51.8753.0 W

The local coding agent, solved. Coding is the workload where generation speed matters most: agent loops burn thousands of tokens on edits, diffs and retries, and a slow model makes the whole loop unusable. That's why this model is significant: 30B-class code quality at 300+ tok/s means a local coding agent that keeps up with you. If I were building a local-first coding setup on a 24GB card today, this is the model I'd build it around.

The efficiency story is the headline. 318 tok/s at 110.5W is 2.87 tokens per watt, the best big-model efficiency figure we've measured. The H200 and H100 sit just behind (298 and 296 tok/s), and even the A10G holds 145 tok/s, which still beats every dense 32B result in our data on any GPU. The ~20GB floor puts it on 24GB consumer cards; pair it with the smaller Qwen2.5-Coder 7B as a fast autocomplete model and you have a full local coding stack on one card.

About Qwen3-Coder 30B A3B. Qwen3-Coder 30B A3B: from Qwen, 31B parameters, on Hugging Face since July 2025, Apache 2.0 licence. 9,073,846 downloads in the last 30 days and 4 community quantizations.

How it compares. H100 80GB HBM3: Qwen3-Coder 30B A3B 296.2 tok/s, Qwen3 30B A3B 283.8 (31B), Qwen3 30B A3B Instruct 2507 299.8 (31B), GLM-4.7-Flash 186.1 (31B), Gemma 4 31B 76.59 (31B). 1 of 4 beat Qwen3-Coder 30B A3B here.

Cost on a rented GPU. 1M generated tokens of Qwen3-Coder 30B A3B: $0.15 on a RTX 3090 ($0.12/hr, 76 min), $0.30 on a RTX 5090 ($0.39/hr, 46 min, 1.9x the cost).

Qwen3-Coder 30B A3B: cost per 1M generated tokens on rented GPUs

NVIDIA GeForce RTX 3090$0.12/hr
NVIDIA GeForce RTX 5090$0.39/hr
NVIDIA GeForce RTX 4090$0.34/hr
NVIDIA A100 40GB SXM4$0.47/hr
NVIDIA RTX PRO 6000 Blackwell Workstation Edition$1.08/hr
NVIDIA L40S$0.79/hr
NVIDIA L4$0.44/hr
NVIDIA A100 80GB SXM4$0.95/hr
NVIDIA H100 80GB HBM3$2.14/hr
NVIDIA H200$3.59/hr
NVIDIA B200$5.98/hr
NVIDIA B300$6.94/hr
GPUCheapest rateSpeed (tok/s)Cost per 1M generated tokens
NVIDIA GeForce RTX 3090$0.12/hr219.1$0.15
NVIDIA GeForce RTX 5090$0.39/hr365.8$0.30
NVIDIA GeForce RTX 4090$0.34/hr271$0.34
NVIDIA A100 40GB SXM4$0.47/hr173.5$0.76
NVIDIA RTX PRO 6000 Blackwell Workstation Edition$1.08/hr317.7$0.94
NVIDIA L40S$0.79/hr219$1.00
NVIDIA L4$0.44/hr99.14$1.23
NVIDIA A100 80GB SXM4$0.95/hr182.3$1.44
NVIDIA H100 80GB HBM3$2.14/hr296.2$2.00
NVIDIA H200$3.59/hr297.9$3.35
NVIDIA B200$5.98/hr277.7$5.98
NVIDIA B300$6.94/hr284.9$6.77

Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.

Speed tiers for Qwen3-Coder 30B A3B. 30+ tok/s: 13 (RTX 5090, RTX 4090, RTX 3090). 30 tok/s is roughly where replies outpace reading.

Reading your prompt. Before Qwen3-Coder 30B A3B writes anything it reads the input: 11452.4 tok/s on the RTX 5090 (0.3s for a 4,000-token prompt), 9315.9 on the RTX 4090 (0.4s), 2424.5 on the L4 (1.6s). Long documents and big code files feel this number more than the generation speed.

VRAM for Qwen3-Coder 30B A3B. Measured peak 17.8GB, so 24GB is the smallest common card size; smallest card it ran on: RTX 4090 (24GB). With long context: Q4_K_M 20GB (tested), Q2_K 13GB, Q3_K_M 16GB, Q5_K_M 24GB, Q6_K 28GB.

Power on Qwen3-Coder 30B A3B. Most efficient: A100 40GB SXM4, 92W, 0.15 kWh per 1M generated tokens. Hungriest: B200, 282W, 0.28 kWh. At $0.15/kWh: $0.022 per 1M generated tokens.

Our verdict

Qwen3 Coder 30B-A3B: 318 tok/s at 110W, the fastest big model we've benchmarked, full stop. MoE speed makes local coding agents actually viable, and the ~20GB floor fits any 24GB card. Our measurements, our harness, logged power on every run.

FAQ

What GPU do I need for Qwen3 Coder 30B-A3B?
A 24GB card. Measured peak was ~20GB at Q4_K_M. RX 7900 XTX, RTX 3090/3090 Ti, RTX 4090, or any datacenter card from the A10G up (145 tok/s there).
Why is it faster than models half its size?
Mixture-of-experts: only ~3B of the 30B parameters fire per token. In our data it out-generates the dense Qwen2.5-Coder 32B by almost 4× (318 vs 83 tok/s) in the same VRAM class.
Is it good enough to replace a cloud coding assistant?
Speed-wise, yes, 300+ tok/s is faster than most API endpoints deliver. Quality-wise it's the strongest local coding model we've benchmarked in this class; whether it replaces a frontier API model depends on your codebase's difficulty.
Qwen3 Coder 30B-A3B or Qwen2.5-Coder 32B?
The MoE, in almost every case: nearly 4× the generation speed in the same memory footprint. The dense 32B is the pick only if its outputs win on your specific code, for agent loops, speed compounds.
What's the best budget setup for it?
A used RTX 3090 (24GB) is the cheapest card that clears the floor with context headroom. For teams, one RTX PRO 6000 serves it at 318 tok/s and 110W, quiet, efficient, office-friendly.