Qwen3-Coder 30B A3B · 10 GPUs measured first-party · llama.cpp Q4_K_M · Updated July 2026
Qwen3 Coder 30B-A3B is the coding-tuned version of the 30B MoE, and the fastest 'big' model in our entire database: 318 tok/s on the RTX PRO 6000 Blackwell at just 110W. Same mixture-of-experts economics as its sibling (30B in memory, ~3B active per token), same ~20GB measured floor, measured on 10 GPUs with our pinned llama.cpp harness.
Benchmarked weights: bartowski/Qwen_Qwen3-Coder-30B-A3B-Instruct-GGUF
What GPU Do You Need for Qwen3-Coder 30B A3B?, tok/s, fastest 10
Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.
Qwen3-Coder 30B A3B. Measured generation speed by GPU
| GPU | tok/s | Prompt t/s | tok/W | Avg power |
|---|---|---|---|---|
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 317.68 | 9786.1 | 2.87 | 110.5 W |
| NVIDIA H200 | 297.9 | 7287.6 | 2.39 | 124.4 W |
| NVIDIA H100 80GB HBM3 | 296.22 | 7235.9 | 2.33 | 127.4 W |
| NVIDIA B300 | 284.88 | 4674.1 | 1.11 | 256.2 W |
| NVIDIA B200 | 277.68 | 7658.5 | 0.98 | 282.0 W |
| NVIDIA L40S | 219.47 | 8071.5 | 1.94 | 113.3 W |
| NVIDIA A100 80GB SXM4 | 182.26 | 3754.5 | 1.67 | 109.0 W |
| NVIDIA A100 40GB SXM4 | 176.74 | 3677.2 | 2.5 | 70.8 W |
| NVIDIA A10G | 145 | 2766 | 1.81 | 80.1 W |
| NVIDIA L4 | 98.84 | 2398.2 | 2.5 | 39.6 W |
The local coding agent, solved. Coding is the workload where generation speed matters most: agent loops burn thousands of tokens on edits, diffs and retries, and a slow model makes the whole loop unusable. That's why this model is significant: 30B-class code quality at 300+ tok/s means a local coding agent that keeps up with you. If I were building a local-first coding setup on a 24GB card today, this is the model I'd build it around.
The efficiency story is the headline. 318 tok/s at 110.5W is 2.87 tokens per watt, the best big-model efficiency figure we've measured. The H200 and H100 sit just behind (298 and 296 tok/s), and even the A10G holds 145 tok/s, which still beats every dense 32B result in our data on any GPU. The ~20GB floor puts it on 24GB consumer cards; pair it with the smaller Qwen2.5-Coder 7B as a fast autocomplete model and you have a full local coding stack on one card.
Qwen3 Coder 30B-A3B: 318 tok/s at 110W, the fastest big model we've benchmarked, full stop. MoE speed makes local coding agents actually viable, and the ~20GB floor fits any 24GB card. Our measurements, our harness, logged power on every run.