Qwen2.5-Coder 14B · 39 cards measured first-party · Updated July 2026

How Fast Does Qwen2.5-Coder 14B Run on Each GPU?

Qwen2.5-Coder 14B is the practical local coding assistant: big enough to be genuinely useful, small enough at ~11.5GB to fit on hardware people actually own. If you want a coding model running on your own machine, this is the benchmark that matters.

Benchmarked weights: Qwen/Qwen2.5-Coder-14B-Instruct-GGUF

Fastest we measured
NVIDIA H100 NVL

NVIDIA H100 NVL

170.3 tok/s on Qwen2.5-Coder 14B. Anchored estimate. 94GB of VRAM, 400W board rating. AI Score 67.0/100 across our full 12-workload suite.

Pros
  • 170.3 tok/s on Qwen2.5-Coder 14B
  • 94GB, clears the Qwen2.5-Coder 14B floor
  • Rentable by the hour rather than bought
Cons
  • 400W board rating
  • Datacenter or workstation hardware, not a retail purchase

Best for: Qwen2.5-Coder 14B work where you want the ceiling gone rather than the cheapest entry.

Runner-up
NVIDIA B300

NVIDIA B300

158.49 tok/s on Qwen2.5-Coder 14B. Measured on our bench. 288GB of VRAM, 1400W board rating. AI Score 93.8/100 across our full 12-workload suite.

Pros
  • 158.49 tok/s on Qwen2.5-Coder 14B
  • 288GB, clears the Qwen2.5-Coder 14B floor
  • Rentable by the hour rather than bought
Cons
  • 1400W board rating
  • Datacenter or workstation hardware, not a retail purchase

Best for: Qwen2.5-Coder 14B work where you want the ceiling gone rather than the cheapest entry.

Third
NVIDIA B200

NVIDIA B200

150.96 tok/s on Qwen2.5-Coder 14B. Measured on our bench. 192GB of VRAM, 1000W board rating. AI Score 78.0/100 across our full 12-workload suite.

Pros
  • 150.96 tok/s on Qwen2.5-Coder 14B
  • 192GB, clears the Qwen2.5-Coder 14B floor
  • Rentable by the hour rather than bought
Cons
  • 1000W board rating
  • Datacenter or workstation hardware, not a retail purchase

Best for: Qwen2.5-Coder 14B work where you want the ceiling gone rather than the cheapest entry.

170.3tok/s
Fastest: NVIDIA H100 NVL
anchored estimate
51
Cards that run Qwen2.5-Coder 14B
of 61 we have data for
10
Cards that can't run it at all
published as hard gates, not omissions
729%
Fastest vs slowest that fits
170.3 vs 23.35 tok/s

Bandwidth-bound like the rest of the LLM ladder. What makes 14B interesting is that it's the point where the whole consumer market is still in play, so the ranking is a clean read on memory bandwidth across every tier, from datacenter HBM down to a mid-range GDDR card.

Qwen2.5-Coder 14B, the 12 fastest cards we have data for

NVIDIA H100 NVL
170.3 tok/s
NVIDIA B300
158.49 tok/s
NVIDIA GH200 Grace Hopper
151.5 tok/s
NVIDIA B200
150.96 tok/s
NVIDIA GeForce RTX 5090
149.75 tok/s
NVIDIA H200
148.43 tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
147.74 tok/s
NVIDIA H100 80GB HBM3
144.84 tok/s
NVIDIA H800 80GB
144.8 tok/s
NVIDIA B100
143.4 tok/s
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
140.4 tok/s
NVIDIA RTX PRO 6000 Blackwell Server Edition
135.07 tok/s

Single stream, batch size 1. 39 of the 61 cards on this page were measured first-party by us; the rest are anchored estimates against those measurements and are labelled in the table below.

Won't fit, Qwen2.5-Coder 14B gates these cards outright

NVIDIA GeForce RTX 308010GB
NVIDIA GeForce RTX 2060 Super8GB
NVIDIA GeForce RTX 2070 SUPER8GB
NVIDIA GeForce RTX 20708GB
NVIDIA GeForce RTX 2080 Super8GB
NVIDIA GeForce RTX 2080 Founders Edition8GB
NVIDIA GeForce RTX 3060 Ti8GB
NVIDIA GeForce RTX 3070 Founders Edition8GB
GeForce RTX 40608GB
NVIDIA GeForce RTX 20606GB
GPUVRAMWhy it fails
NVIDIA GeForce RTX 308010GBrequires ~11.5GB VRAM
NVIDIA GeForce RTX 2060 Super8GBNeeds needs ~10GB VRAM
NVIDIA GeForce RTX 2070 SUPER8GBNeeds needs ~10GB VRAM
NVIDIA GeForce RTX 20708GBNeeds needs ~10GB VRAM
NVIDIA GeForce RTX 2080 Super8GBNeeds needs ~10GB VRAM
NVIDIA GeForce RTX 2080 Founders Edition8GBNeeds needs ~10GB VRAM
NVIDIA GeForce RTX 3060 Ti8GBrequires ~11.5GB VRAM
NVIDIA GeForce RTX 3070 Founders Edition8GBrequires ~11.5GB VRAM
GeForce RTX 40608GBrequires ~11.5GB VRAM
NVIDIA GeForce RTX 20606GBNeeds needs ~10GB VRAM

No driver update fixes a VRAM ceiling.

Full Qwen2.5-Coder 14B leaderboard, every card that runs it

NVIDIA H100 NVL170.3 tok/s
NVIDIA B300158.49 tok/s
NVIDIA GH200 Grace Hopper151.5 tok/s
NVIDIA B200150.96 tok/s
NVIDIA GeForce RTX 5090149.75 tok/s
NVIDIA H200148.43 tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition147.74 tok/s
NVIDIA H100 80GB HBM3144.84 tok/s
NVIDIA H800 80GB144.8 tok/s
NVIDIA B100143.4 tok/s
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition140.4 tok/s
NVIDIA RTX PRO 6000 Blackwell Server Edition135.07 tok/s
NVIDIA RTX PRO 5000 Blackwell114.63 tok/s
NVIDIA GeForce RTX 409095.12 tok/s
NVIDIA A800 80GB89.5 tok/s
NVIDIA A100 80GB SXM489.45 tok/s
NVIDIA A100 80GB PCIe88.36 tok/s
NVIDIA RTX 6000 Ada Generation87.51 tok/s
NVIDIA H100 PCIe86.5 tok/s
GeForce RTX 5070 Ti83.44 tok/s
GeForce RTX 508081.97 tok/s
NVIDIA RTX PRO 4500 Blackwell80.09 tok/s
NVIDIA GeForce RTX 309078.79 tok/s
NVIDIA GeForce RTX 3080 Ti78.31 tok/s
NVIDIA L40S74.33 tok/s
NVIDIA L4074.27 tok/s
NVIDIA A100 40GB PCIe71.0 tok/s
NVIDIA GeForce RTX 408069.93 tok/s
NVIDIA RTX A550069.9 tok/s
NVIDIA RTX A600068.51 tok/s
NVIDIA A100 40GB SXM468.2 tok/s
AMD Radeon Pro W790066.6 tok/s
NVIDIA RTX A500064.89 tok/s
NVIDIA RTX 5880 Ada Generation63.4 tok/s
NVIDIA RTX PRO 4000 Blackwell60.62 tok/s
NVIDIA RTX 5000 Ada Generation56.66 tok/s
NVIDIA RTX A450054.29 tok/s
NVIDIA GeForce RTX 407050.74 tok/s
NVIDIA A10G46.96 tok/s
NVIDIA Quadro RTX 800046.74 tok/s
NVIDIA Quadro RTX 6000 (Turing)46.54 tok/s
NVIDIA Quadro RTX 500045.1 tok/s
NVIDIA RTX 4500 Ada Generation43.7 tok/s
AMD Radeon Pro W680043.6 tok/s
AMD Radeon RX 6900 XT43.6 tok/s
NVIDIA RTX A400041.13 tok/s
NVIDIA RTX 4000 (Ada Generation)36.34 tok/s
NVIDIA GeForce RTX 306035.52 tok/s
NVIDIA GeForce RTX 4060 Ti31.12 tok/s
NVIDIA L427.46 tok/s
NVIDIA RTX 2000 Ada Generation23.35 tok/s
GPUResultVRAMSource
NVIDIA H100 NVL170.3 tok/s94GBEstimated
NVIDIA B300158.49 tok/s288GBMeasured
NVIDIA GH200 Grace Hopper151.5 tok/s141GBEstimated
NVIDIA B200150.96 tok/s192GBMeasured
NVIDIA GeForce RTX 5090149.75 tok/s32GBMeasured
NVIDIA H200148.43 tok/s141GBMeasured
NVIDIA RTX PRO 6000 Blackwell Workstation Edition147.74 tok/s96GBMeasured
NVIDIA H100 80GB HBM3144.84 tok/s80GBMeasured
NVIDIA H800 80GB144.8 tok/s80GBEstimated
NVIDIA B100143.4 tok/s192GBEstimated
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition140.4 tok/s96GBEstimated
NVIDIA RTX PRO 6000 Blackwell Server Edition135.07 tok/s96GBMeasured
NVIDIA RTX PRO 5000 Blackwell114.63 tok/s48GBMeasured
NVIDIA GeForce RTX 409095.12 tok/s24GBMeasured
NVIDIA A800 80GB89.5 tok/s80GBEstimated
NVIDIA A100 80GB SXM489.45 tok/s80GBMeasured
NVIDIA A100 80GB PCIe88.36 tok/s80GBMeasured
NVIDIA RTX 6000 Ada Generation87.51 tok/s48GBMeasured
NVIDIA H100 PCIe86.5 tok/s80GBEstimated
GeForce RTX 5070 Ti83.44 tok/s16GBMeasured
GeForce RTX 508081.97 tok/s16GBMeasured
NVIDIA RTX PRO 4500 Blackwell80.09 tok/s32GBMeasured
NVIDIA GeForce RTX 309078.79 tok/s24GBMeasured
NVIDIA GeForce RTX 3080 Ti78.31 tok/s12GBMeasured
NVIDIA L40S74.33 tok/s48GBMeasured
NVIDIA L4074.27 tok/s48GBMeasured
NVIDIA A100 40GB PCIe71.0 tok/s40GBEstimated
NVIDIA GeForce RTX 408069.93 tok/s16GBMeasured
NVIDIA RTX A550069.9 tok/s24GBEstimated
NVIDIA RTX A600068.51 tok/s48GBMeasured
NVIDIA A100 40GB SXM468.2 tok/s40GBEstimated
AMD Radeon Pro W790066.6 tok/s48GBEstimated
NVIDIA RTX A500064.89 tok/s24GBMeasured
NVIDIA RTX 5880 Ada Generation63.4 tok/s48GBEstimated
NVIDIA RTX PRO 4000 Blackwell60.62 tok/s24GBMeasured
NVIDIA RTX 5000 Ada Generation56.66 tok/s32GBMeasured
NVIDIA RTX A450054.29 tok/s20GBMeasured
NVIDIA GeForce RTX 407050.74 tok/s12GBMeasured
NVIDIA A10G46.96 tok/s24GBMeasured
NVIDIA Quadro RTX 800046.74 tok/s48GBMeasured
NVIDIA Quadro RTX 6000 (Turing)46.54 tok/s24GBMeasured
NVIDIA Quadro RTX 500045.1 tok/s16GBEstimated
NVIDIA RTX 4500 Ada Generation43.7 tok/s24GBEstimated
AMD Radeon Pro W680043.6 tok/s32GBEstimated
AMD Radeon RX 6900 XT43.6 tok/s16GBEstimated
NVIDIA RTX A400041.13 tok/s16GBMeasured
NVIDIA RTX 4000 (Ada Generation)36.34 tok/s20GBMeasured
NVIDIA GeForce RTX 306035.52 tok/s12GBMeasured
NVIDIA GeForce RTX 4060 Ti31.12 tok/s16GBMeasured
NVIDIA L427.46 tok/s24GBMeasured
NVIDIA RTX 2000 Ada Generation23.35 tok/s16GBMeasured

Tap any column to sort. Measured = we rented and ran this card ourselves. Estimated = interpolated against our measured anchors, never blended silently.

Because this workload is bandwidth-bound, the ranking above tracks memory bandwidth far more closely than core counts or price. A card with fewer tensor cores and faster memory will beat a card with the opposite.

That's the reason we run twelve workloads instead of publishing one score. A GPU isn't fast or slow. It's fast at some things and gated out of others, and which of those matters depends entirely on what you're actually going to run.

Our verdict

NVIDIA H100 NVL tops our Qwen2.5-Coder 14B leaderboard at 170.3 tok/s (anchored estimate), 729% of the way clear of the slowest card that still fits. But the number that decides most purchases isn't on the chart. It's the 10 cards that can't run Qwen2.5-Coder 14B at all. This is a bandwidth workload: buy memory speed, not tensor cores.

FAQ

What is the fastest GPU for Qwen2.5-Coder 14B?
NVIDIA H100 NVL, at 170.3 tok/s on our bench, an anchored estimate against our measured cards. It carries 94GB of VRAM. Of the 61 cards we have Qwen2.5-Coder 14B data for, 51 can run it at all.
How much VRAM do I need for Qwen2.5-Coder 14B?
~11.5GB at Q4_K_M, comfortable on a 12GB card, easy on 16GB.
Why does the Qwen2.5-Coder 14B ranking look different from your other benchmarks?
Because this workload is bandwidth-bound, the ranking above tracks memory bandwidth far more closely than core counts or price. A card with fewer tensor cores and faster memory will beat a card with the opposite. That's why we publish twelve separate workloads rather than one blended score, the ordering genuinely changes depending on the job.
Are these Qwen2.5-Coder 14B numbers measured or estimated?
Both, and every row says which. 39 of the 61 cards here were rented and run by us on the same harness. The remainder are anchored estimates interpolated per workload against those measurements. We never blend the two silently, if a row says Estimated, we have not run that card.
Can I rent a GPU to run Qwen2.5-Coder 14B instead of buying one?
Yes, and for the cards at the top of this leaderboard it's the only realistic option, most of them have no retail channel at all. It's also how we got these numbers: we rented the hardware by the hour rather than buying it. That's worth considering before you spend on a card to find out whether it's fast enough.
Why publish cards that can't run Qwen2.5-Coder 14B?
Because it's the most useful thing we know. A card that can't load a model doesn't run it slowly, it doesn't run it. Most benchmark sites leave that as a blank cell or quietly drop to a smaller quantisation to produce a number. We publish it as a hard gate and score it zero, because 'this card cannot do the thing you want' is the answer to the question you were actually asking.

How we test

Every ranking on this page comes from our own benchmark runs, not vendor claims. Cards marked Measured were rented and run by us; cards marked Estimated are interpolated per workload against those measured anchors and are labelled on every row, we never blend the two silently. LLMs run on llama.cpp (llama-bench) at Q4_K_M with -p 512 -n 128. Diffusion and video run on diffusers/ComfyUI at BF16, with SDXL at FP16. Each workload gets a warmup pass plus multiple timed runs (5 for small LLMs, 3 for large models and images, 2 for video); we publish the mean as the result and the minimum as the 1% low. Run-to-run variance is under 0.5%. Telemetry, power, temperature, utilisation, clocks, peak VRAM, is sampled at 1 Hz for the duration of every run. Where a model exceeds a card's VRAM we publish a hard won't-fit result rather than quietly dropping to a smaller quantisation. A card that can't run a model scores zero on it. Silently swapping precision to make a number appear would make every number on this site meaningless. All figures are single-GPU, single-stream, batch-size-1. That is the honest way to measure what one card does for one user, and it is deliberately not how a datacenter serves a model. Vendor and MLPerf figures use large batches across many GPUs and will be far higher. Neither is wrong, they answer different questions. Ours answers 'what will this card do for me'.