Qwen2.5-Coder 14B · 39 cards measured first-party · Updated July 2026
Qwen2.5-Coder 14B is the practical local coding assistant: big enough to be genuinely useful, small enough at ~11.5GB to fit on hardware people actually own. If you want a coding model running on your own machine, this is the benchmark that matters.
Benchmarked weights: Qwen/Qwen2.5-Coder-14B-Instruct-GGUF

170.3 tok/s on Qwen2.5-Coder 14B. Anchored estimate. 94GB of VRAM, 400W board rating. AI Score 67.0/100 across our full 12-workload suite.
Best for: Qwen2.5-Coder 14B work where you want the ceiling gone rather than the cheapest entry.

158.49 tok/s on Qwen2.5-Coder 14B. Measured on our bench. 288GB of VRAM, 1400W board rating. AI Score 93.8/100 across our full 12-workload suite.
Best for: Qwen2.5-Coder 14B work where you want the ceiling gone rather than the cheapest entry.

150.96 tok/s on Qwen2.5-Coder 14B. Measured on our bench. 192GB of VRAM, 1000W board rating. AI Score 78.0/100 across our full 12-workload suite.
Best for: Qwen2.5-Coder 14B work where you want the ceiling gone rather than the cheapest entry.
Bandwidth-bound like the rest of the LLM ladder. What makes 14B interesting is that it's the point where the whole consumer market is still in play, so the ranking is a clean read on memory bandwidth across every tier, from datacenter HBM down to a mid-range GDDR card.
Qwen2.5-Coder 14B, the 12 fastest cards we have data for
Single stream, batch size 1. 39 of the 61 cards on this page were measured first-party by us; the rest are anchored estimates against those measurements and are labelled in the table below.
Won't fit, Qwen2.5-Coder 14B gates these cards outright
| GPU | VRAM | Why it fails |
|---|---|---|
| NVIDIA GeForce RTX 3080 | 10GB | requires ~11.5GB VRAM |
| NVIDIA GeForce RTX 2060 Super | 8GB | Needs needs ~10GB VRAM |
| NVIDIA GeForce RTX 2070 SUPER | 8GB | Needs needs ~10GB VRAM |
| NVIDIA GeForce RTX 2070 | 8GB | Needs needs ~10GB VRAM |
| NVIDIA GeForce RTX 2080 Super | 8GB | Needs needs ~10GB VRAM |
| NVIDIA GeForce RTX 2080 Founders Edition | 8GB | Needs needs ~10GB VRAM |
| NVIDIA GeForce RTX 3060 Ti | 8GB | requires ~11.5GB VRAM |
| NVIDIA GeForce RTX 3070 Founders Edition | 8GB | requires ~11.5GB VRAM |
| GeForce RTX 4060 | 8GB | requires ~11.5GB VRAM |
| NVIDIA GeForce RTX 2060 | 6GB | Needs needs ~10GB VRAM |
No driver update fixes a VRAM ceiling.
Full Qwen2.5-Coder 14B leaderboard, every card that runs it
| GPU | Result | VRAM | Source |
|---|---|---|---|
| NVIDIA H100 NVL | 170.3 tok/s | 94GB | Estimated |
| NVIDIA B300 | 158.49 tok/s | 288GB | Measured |
| NVIDIA GH200 Grace Hopper | 151.5 tok/s | 141GB | Estimated |
| NVIDIA B200 | 150.96 tok/s | 192GB | Measured |
| NVIDIA GeForce RTX 5090 | 149.75 tok/s | 32GB | Measured |
| NVIDIA H200 | 148.43 tok/s | 141GB | Measured |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 147.74 tok/s | 96GB | Measured |
| NVIDIA H100 80GB HBM3 | 144.84 tok/s | 80GB | Measured |
| NVIDIA H800 80GB | 144.8 tok/s | 80GB | Estimated |
| NVIDIA B100 | 143.4 tok/s | 192GB | Estimated |
| NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition | 140.4 tok/s | 96GB | Estimated |
| NVIDIA RTX PRO 6000 Blackwell Server Edition | 135.07 tok/s | 96GB | Measured |
| NVIDIA RTX PRO 5000 Blackwell | 114.63 tok/s | 48GB | Measured |
| NVIDIA GeForce RTX 4090 | 95.12 tok/s | 24GB | Measured |
| NVIDIA A800 80GB | 89.5 tok/s | 80GB | Estimated |
| NVIDIA A100 80GB SXM4 | 89.45 tok/s | 80GB | Measured |
| NVIDIA A100 80GB PCIe | 88.36 tok/s | 80GB | Measured |
| NVIDIA RTX 6000 Ada Generation | 87.51 tok/s | 48GB | Measured |
| NVIDIA H100 PCIe | 86.5 tok/s | 80GB | Estimated |
| GeForce RTX 5070 Ti | 83.44 tok/s | 16GB | Measured |
| GeForce RTX 5080 | 81.97 tok/s | 16GB | Measured |
| NVIDIA RTX PRO 4500 Blackwell | 80.09 tok/s | 32GB | Measured |
| NVIDIA GeForce RTX 3090 | 78.79 tok/s | 24GB | Measured |
| NVIDIA GeForce RTX 3080 Ti | 78.31 tok/s | 12GB | Measured |
| NVIDIA L40S | 74.33 tok/s | 48GB | Measured |
| NVIDIA L40 | 74.27 tok/s | 48GB | Measured |
| NVIDIA A100 40GB PCIe | 71.0 tok/s | 40GB | Estimated |
| NVIDIA GeForce RTX 4080 | 69.93 tok/s | 16GB | Measured |
| NVIDIA RTX A5500 | 69.9 tok/s | 24GB | Estimated |
| NVIDIA RTX A6000 | 68.51 tok/s | 48GB | Measured |
| NVIDIA A100 40GB SXM4 | 68.2 tok/s | 40GB | Estimated |
| AMD Radeon Pro W7900 | 66.6 tok/s | 48GB | Estimated |
| NVIDIA RTX A5000 | 64.89 tok/s | 24GB | Measured |
| NVIDIA RTX 5880 Ada Generation | 63.4 tok/s | 48GB | Estimated |
| NVIDIA RTX PRO 4000 Blackwell | 60.62 tok/s | 24GB | Measured |
| NVIDIA RTX 5000 Ada Generation | 56.66 tok/s | 32GB | Measured |
| NVIDIA RTX A4500 | 54.29 tok/s | 20GB | Measured |
| NVIDIA GeForce RTX 4070 | 50.74 tok/s | 12GB | Measured |
| NVIDIA A10G | 46.96 tok/s | 24GB | Measured |
| NVIDIA Quadro RTX 8000 | 46.74 tok/s | 48GB | Measured |
| NVIDIA Quadro RTX 6000 (Turing) | 46.54 tok/s | 24GB | Measured |
| NVIDIA Quadro RTX 5000 | 45.1 tok/s | 16GB | Estimated |
| NVIDIA RTX 4500 Ada Generation | 43.7 tok/s | 24GB | Estimated |
| AMD Radeon Pro W6800 | 43.6 tok/s | 32GB | Estimated |
| AMD Radeon RX 6900 XT | 43.6 tok/s | 16GB | Estimated |
| NVIDIA RTX A4000 | 41.13 tok/s | 16GB | Measured |
| NVIDIA RTX 4000 (Ada Generation) | 36.34 tok/s | 20GB | Measured |
| NVIDIA GeForce RTX 3060 | 35.52 tok/s | 12GB | Measured |
| NVIDIA GeForce RTX 4060 Ti | 31.12 tok/s | 16GB | Measured |
| NVIDIA L4 | 27.46 tok/s | 24GB | Measured |
| NVIDIA RTX 2000 Ada Generation | 23.35 tok/s | 16GB | Measured |
Tap any column to sort. Measured = we rented and ran this card ourselves. Estimated = interpolated against our measured anchors, never blended silently.
Because this workload is bandwidth-bound, the ranking above tracks memory bandwidth far more closely than core counts or price. A card with fewer tensor cores and faster memory will beat a card with the opposite.
That's the reason we run twelve workloads instead of publishing one score. A GPU isn't fast or slow. It's fast at some things and gated out of others, and which of those matters depends entirely on what you're actually going to run.
NVIDIA H100 NVL tops our Qwen2.5-Coder 14B leaderboard at 170.3 tok/s (anchored estimate), 729% of the way clear of the slowest card that still fits. But the number that decides most purchases isn't on the chart. It's the 10 cards that can't run Qwen2.5-Coder 14B at all. This is a bandwidth workload: buy memory speed, not tensor cores.
Every ranking on this page comes from our own benchmark runs, not vendor claims. Cards marked Measured were rented and run by us; cards marked Estimated are interpolated per workload against those measured anchors and are labelled on every row, we never blend the two silently. LLMs run on llama.cpp (llama-bench) at Q4_K_M with -p 512 -n 128. Diffusion and video run on diffusers/ComfyUI at BF16, with SDXL at FP16. Each workload gets a warmup pass plus multiple timed runs (5 for small LLMs, 3 for large models and images, 2 for video); we publish the mean as the result and the minimum as the 1% low. Run-to-run variance is under 0.5%. Telemetry, power, temperature, utilisation, clocks, peak VRAM, is sampled at 1 Hz for the duration of every run. Where a model exceeds a card's VRAM we publish a hard won't-fit result rather than quietly dropping to a smaller quantisation. A card that can't run a model scores zero on it. Silently swapping precision to make a number appear would make every number on this site meaningless. All figures are single-GPU, single-stream, batch-size-1. That is the honest way to measure what one card does for one user, and it is deliberately not how a datacenter serves a model. Vendor and MLPerf figures use large batches across many GPUs and will be far higher. Neither is wrong, they answer different questions. Ours answers 'what will this card do for me'.