Qwen2.5-Coder 14B · 62 cards measured first-party · Updated October 2026

How Fast Does Qwen2.5-Coder 14B Run on Each GPU?

Qwen2.5-Coder 14B is the practical local coding assistant: big enough to be genuinely useful, small enough at ~11.5GB to fit on hardware people actually own. If you want a coding model running on your own machine, this is the benchmark that matters.

Benchmarked weights: Qwen/Qwen2.5-Coder-14B-Instruct-GGUF

Fastest we measured
NVIDIA B300

NVIDIA B300

158.5 tok/s on Qwen2.5-Coder 14B, the ceiling. Measured on our bench. 288GB of VRAM, $40,000 at launch.

Pros
  • 158.5 tok/s on Qwen2.5-Coder 14B
  • 288GB, clears the Qwen2.5-Coder 14B floor
  • Rentable by the hour rather than bought
Cons
  • 1400W board rating
  • Datacenter or workstation hardware, not a retail purchase
Best consumer card
NVIDIA GeForce RTX 5090

NVIDIA GeForce RTX 5090

149.8 tok/s on Qwen2.5-Coder 14B, fastest card you can buy at retail. Measured on our bench. 32GB of VRAM, $1,999 at launch.

Pros
  • 149.8 tok/s on Qwen2.5-Coder 14B
  • 32GB, clears the Qwen2.5-Coder 14B floor
Cons
  • 575W board rating
Cheapest card that runs it
Intel Arc B580

Intel Arc B580

33.0 tok/s on Qwen2.5-Coder 14B, lowest launch price that still fits. Anchored estimate. 12GB of VRAM, $179 at launch.

Pros
  • 33.0 tok/s on Qwen2.5-Coder 14B
  • 12GB, clears the Qwen2.5-Coder 14B floor
Cons
  • 190W board rating
Best value
GeForce RTX 5070 Ti

GeForce RTX 5070 Ti

83.44 tok/s on Qwen2.5-Coder 14B, most speed per dollar. Measured on our bench. 16GB of VRAM, $749 at launch. That is 111.4 tok/s per $1,000 of launch price.

Pros
  • 83.44 tok/s on Qwen2.5-Coder 14B
  • 16GB, clears the Qwen2.5-Coder 14B floor
Cons
  • 300W board rating
158.5tok/s
Fastest: NVIDIA B300
measured
78
Cards that run Qwen2.5-Coder 14B
of 102 we have data for
24
Cards that can't run it at all
published as hard gates, not omissions
757%
Fastest vs slowest that fits
158.5 vs 18.5 tok/s

Bandwidth-bound like the rest of the LLM ladder. What makes 14B interesting is that it's the point where the whole consumer market is still in play, so the ranking is a clean read on memory bandwidth across every tier, from datacenter HBM down to a mid-range GDDR card.

Qwen2.5-Coder 14B: speed on every GPU we have data for

NVIDIA B300
158.5 tok/s
NVIDIA GH200 Grace Hopper
151.5 tok/s
NVIDIA B200
151 tok/s
NVIDIA GeForce RTX 5090
149.8 tok/s
NVIDIA H200
148.4 tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
147.7 tok/s
NVIDIA H100 80GB HBM3
144.8 tok/s
NVIDIA H800 80GB
144.8 tok/s
NVIDIA B100
143.4 tok/s
NVIDIA RTX PRO 6000 Blackwell Server Edition
135.1 tok/s
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
125.8 tok/s
NVIDIA H100 NVL
125.6 tok/s
NVIDIA RTX PRO 5000 Blackwell
114.6 tok/s
NVIDIA H100 PCIe
103.7 tok/s
NVIDIA GeForce RTX 4090
95.12 tok/s

Top 15 shown; 63 more cards in the full table below.

Single stream, batch size 1. 39 of the 61 cards on this page were measured first-party by us; the rest are anchored estimates against those measurements and are labelled in the table below.

Efficiency: tok/s per 100W drawn

NVIDIA GeForce RTX 5090
127.88 tok/s / 100W
NVIDIA H200
120.28 tok/s / 100W
NVIDIA RTX PRO 4500 Blackwell
76.28 tok/s / 100W
NVIDIA RTX PRO 5000 Blackwell
63.3 tok/s / 100W
NVIDIA H100 80GB HBM3
62.19 tok/s / 100W
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
61.74 tok/s / 100W
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
60.4 tok/s / 100W
GeForce RTX 5080
56.03 tok/s / 100W
NVIDIA RTX A4000
55.36 tok/s / 100W
NVIDIA GeForce RTX 4090
54.7 tok/s / 100W
NVIDIA H100 PCIe
53.4 tok/s / 100W
NVIDIA RTX PRO 6000 Blackwell Server Edition
51.75 tok/s / 100W
NVIDIA A100 80GB SXM4
51.14 tok/s / 100W
NVIDIA RTX 2000 Ada Generation
50.87 tok/s / 100W
GeForce RTX 4080 Super
50.43 tok/s / 100W

Top 15 shown; 40 more cards in the full table below.

Power is the average pulled during the run, sampled at 1Hz. The fastest card is often not the one here, and for anything left running this is the number that shows up on the bill.

Value: tok/s per $1,000 of MSRP

GeForce RTX 5070 Ti
111.4 tok/s / $1k
NVIDIA GeForce RTX 3060
107.96 tok/s / $1k
GeForce RTX 5070
105.92 tok/s / $1k
GeForce RTX 5060 Ti
105.03 tok/s / $1k
NVIDIA GeForce RTX 4070 Super
85.53 tok/s / $1k
NVIDIA GeForce RTX 4070
84.71 tok/s / $1k
NVIDIA GeForce RTX 4070 Ti Super
82.47 tok/s / $1k
GeForce RTX 5080
82.05 tok/s / $1k
NVIDIA GeForce RTX 5090
74.91 tok/s / $1k
GeForce RTX 4080 Super
70.62 tok/s / $1k
NVIDIA GeForce RTX 3080 Ti
65.31 tok/s / $1k
NVIDIA GeForce RTX 4070 Ti
63.49 tok/s / $1k
NVIDIA GeForce RTX 4060 Ti 16GB
62.36 tok/s / $1k
NVIDIA GeForce RTX 4090
59.49 tok/s / $1k
NVIDIA GeForce RTX 4080
58.32 tok/s / $1k

Top 15 shown; 40 more cards in the full table below.

Launch price, not street price, so it ages. A speed leaderboard always crowns the most expensive card; this is the counterweight.

Won't fit, Qwen2.5-Coder 14B gates these cards outright

NVIDIA GeForce RTX 2080 Ti Founders Edition11GB
AMD Radeon RX 670010GB
NVIDIA GeForce RTX 308010GB
AMD Radeon RX 76008GB
NVIDIA GeForce GTX 1070 Ti8GB
NVIDIA GeForce GTX 10808GB
NVIDIA GeForce RTX 2060 Super8GB
NVIDIA GeForce RTX 2070 SUPER8GB
NVIDIA GeForce RTX 20708GB
NVIDIA GeForce RTX 2080 Super8GB
NVIDIA GeForce RTX 2080 Founders Edition8GB
NVIDIA GeForce RTX 30508GB
NVIDIA GeForce RTX 3060 Ti8GB
NVIDIA GeForce RTX 3070 Ti8GB
NVIDIA GeForce RTX 3070 Founders Edition8GB
GeForce RTX 40608GB
NVIDIA GeForce RTX 50508GB
NVIDIA GeForce RTX 50608GB
Intel Arc A7508GB
NVIDIA GeForce GTX 1660 Super6GB
NVIDIA GeForce GTX 1660 Ti6GB
NVIDIA GeForce GTX 16606GB
NVIDIA GeForce RTX 20606GB
NVIDIA RTX A20006GB
GPUVRAMWhy it fails
NVIDIA GeForce RTX 2080 Ti Founders Edition11GBNeeds ~12GB VRAM
AMD Radeon RX 670010GBNeeds ~12GB VRAM
NVIDIA GeForce RTX 308010GBNeeds ~12GB VRAM
AMD Radeon RX 76008GBNeeds ~10GB VRAM
NVIDIA GeForce GTX 1070 Ti8GBNeeds ~10GB VRAM
NVIDIA GeForce GTX 10808GBNeeds ~10GB VRAM
NVIDIA GeForce RTX 2060 Super8GBNeeds ~10GB VRAM
NVIDIA GeForce RTX 2070 SUPER8GBNeeds ~10GB VRAM
NVIDIA GeForce RTX 20708GBNeeds ~10GB VRAM
NVIDIA GeForce RTX 2080 Super8GBNeeds ~10GB VRAM
NVIDIA GeForce RTX 2080 Founders Edition8GBNeeds ~10GB VRAM
NVIDIA GeForce RTX 30508GBNeeds ~10GB VRAM
NVIDIA GeForce RTX 3060 Ti8GBNeeds ~11.5GB VRAM
NVIDIA GeForce RTX 3070 Ti8GBNeeds ~11.5GB VRAM
NVIDIA GeForce RTX 3070 Founders Edition8GBNeeds ~11.5GB VRAM
GeForce RTX 40608GBNeeds ~11.5GB VRAM
NVIDIA GeForce RTX 50508GBNeeds ~10GB VRAM
NVIDIA GeForce RTX 50608GBNeeds ~10GB VRAM
Intel Arc A7508GBNeeds ~10GB VRAM
NVIDIA GeForce GTX 1660 Super6GBNeeds ~10GB VRAM
NVIDIA GeForce GTX 1660 Ti6GBNeeds ~10GB VRAM
NVIDIA GeForce GTX 16606GBNeeds ~10GB VRAM
NVIDIA GeForce RTX 20606GBNeeds ~10GB VRAM
NVIDIA RTX A20006GBNeeds ~11.5GB VRAM

No driver update fixes a VRAM ceiling.

Full Qwen2.5-Coder 14B leaderboard, every card that runs it

NVIDIA B300158.5 tok/s
NVIDIA GH200 Grace Hopper151.5 tok/s
NVIDIA B200151.0 tok/s
NVIDIA GeForce RTX 5090149.8 tok/s
NVIDIA H200148.4 tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition147.7 tok/s
NVIDIA H100 80GB HBM3144.8 tok/s
NVIDIA H800 80GB144.8 tok/s
NVIDIA B100143.4 tok/s
NVIDIA RTX PRO 6000 Blackwell Server Edition135.1 tok/s
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition125.8 tok/s
NVIDIA H100 NVL125.6 tok/s
NVIDIA RTX PRO 5000 Blackwell114.6 tok/s
NVIDIA H100 PCIe103.7 tok/s
NVIDIA GeForce RTX 409095.12 tok/s
NVIDIA A800 80GB89.5 tok/s
NVIDIA A100 80GB SXM489.45 tok/s
NVIDIA A100 80GB PCIe88.36 tok/s
NVIDIA GeForce RTX 3090 Ti88.25 tok/s
NVIDIA RTX 6000 Ada Generation87.51 tok/s
NVIDIA RTX 5880 Ada Generation86.43 tok/s
NVIDIA A100 40GB SXM486.2 tok/s
NVIDIA A100 40GB PCIe84.11 tok/s
GeForce RTX 5070 Ti83.44 tok/s
GeForce RTX 508081.97 tok/s
AMD Radeon RX 7900 XTX80.3 tok/s
NVIDIA RTX PRO 4500 Blackwell80.09 tok/s
NVIDIA GeForce RTX 309078.79 tok/s
NVIDIA GeForce RTX 3080 Ti78.31 tok/s
NVIDIA L40S74.6 tok/s
NVIDIA L4074.27 tok/s
GeForce RTX 4080 Super70.55 tok/s
NVIDIA GeForce RTX 408069.93 tok/s
NVIDIA RTX A550069.9 tok/s
NVIDIA RTX A600068.51 tok/s
AMD Radeon Pro W790066.6 tok/s
NVIDIA GeForce RTX 4070 Ti Super65.89 tok/s
AMD Radeon RX 7900 XT65.1 tok/s
NVIDIA RTX A500064.89 tok/s
NVIDIA RTX PRO 4000 Blackwell60.62 tok/s
NVIDIA Titan RTX59.75 tok/s
AMD Radeon RX 9070 XT59.0 tok/s
GeForce RTX 507058.15 tok/s
NVIDIA A4057.85 tok/s
NVIDIA RTX 5000 Ada Generation56.66 tok/s
AMD Radeon RX 907056.6 tok/s
NVIDIA TITAN V56.25 tok/s
NVIDIA RTX A450054.29 tok/s
AMD Radeon RX 7800 XT52.8 tok/s
NVIDIA GeForce RTX 4070 Super51.23 tok/s
NVIDIA GeForce RTX 407050.74 tok/s
NVIDIA GeForce RTX 4070 Ti50.73 tok/s
NVIDIA A10G47.07 tok/s
NVIDIA Quadro RTX 800046.74 tok/s
NVIDIA Quadro RTX 6000 (Turing)46.54 tok/s
AMD Radeon RX 6900 XT45.8 tok/s
GeForce RTX 5060 Ti45.06 tok/s
AMD Radeon Pro W780044.1 tok/s
AMD Radeon RX 6950 XT44.1 tok/s
AMD Radeon Pro W680043.6 tok/s
AMD Radeon RX 6800 XT43.6 tok/s
AMD Radeon RX 680043.6 tok/s
NVIDIA RTX 4500 Ada Generation43.38 tok/s
NVIDIA RTX A400041.13 tok/s
NVIDIA Quadro RTX 500041.11 tok/s
AMD Radeon RX 7700 XT37.2 tok/s
NVIDIA RTX 4000 (Ada Generation)36.34 tok/s
NVIDIA TITAN Xp36.2 tok/s
NVIDIA GeForce RTX 306035.52 tok/s
Intel Arc B58033.0 tok/s
NVIDIA GeForce RTX 4060 Ti 16GB31.12 tok/s
NVIDIA TITAN X (Pascal)30.5 tok/s
NVIDIA L427.46 tok/s
GeForce GTX 1080 Ti27.36 tok/s
Intel Arc A770 Limited Edition24.9 tok/s
NVIDIA RTX 2000 Ada Generation23.35 tok/s
NVIDIA T419.76 tok/s
Intel Arc Pro A6018.5 tok/s
GPUResultVRAMSource
NVIDIA B300158.5 tok/s288GBMeasured
NVIDIA GH200 Grace Hopper151.5 tok/s141GBEstimated
NVIDIA B200151.0 tok/s192GBMeasured
NVIDIA GeForce RTX 5090149.8 tok/s32GBMeasured
NVIDIA H200148.4 tok/s141GBMeasured
NVIDIA RTX PRO 6000 Blackwell Workstation Edition147.7 tok/s96GBMeasured
NVIDIA H100 80GB HBM3144.8 tok/s80GBMeasured
NVIDIA H800 80GB144.8 tok/s80GBEstimated
NVIDIA B100143.4 tok/s192GBEstimated
NVIDIA RTX PRO 6000 Blackwell Server Edition135.1 tok/s96GBMeasured
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition125.8 tok/s96GBMeasured
NVIDIA H100 NVL125.6 tok/s94GBMeasured
NVIDIA RTX PRO 5000 Blackwell114.6 tok/s48GBMeasured
NVIDIA H100 PCIe103.7 tok/s80GBMeasured
NVIDIA GeForce RTX 409095.12 tok/s24GBMeasured
NVIDIA A800 80GB89.5 tok/s80GBEstimated
NVIDIA A100 80GB SXM489.45 tok/s80GBMeasured
NVIDIA A100 80GB PCIe88.36 tok/s80GBMeasured
NVIDIA GeForce RTX 3090 Ti88.25 tok/s24GBMeasured
NVIDIA RTX 6000 Ada Generation87.51 tok/s48GBMeasured
NVIDIA RTX 5880 Ada Generation86.43 tok/s48GBMeasured
NVIDIA A100 40GB SXM486.2 tok/s40GBMeasured
NVIDIA A100 40GB PCIe84.11 tok/s40GBMeasured
GeForce RTX 5070 Ti83.44 tok/s16GBMeasured
GeForce RTX 508081.97 tok/s16GBMeasured
AMD Radeon RX 7900 XTX80.3 tok/s24GBEstimated
NVIDIA RTX PRO 4500 Blackwell80.09 tok/s32GBMeasured
NVIDIA GeForce RTX 309078.79 tok/s24GBMeasured
NVIDIA GeForce RTX 3080 Ti78.31 tok/s12GBMeasured
NVIDIA L40S74.6 tok/s48GBMeasured
NVIDIA L4074.27 tok/s48GBMeasured
GeForce RTX 4080 Super70.55 tok/s16GBMeasured
NVIDIA GeForce RTX 408069.93 tok/s16GBMeasured
NVIDIA RTX A550069.9 tok/s24GBEstimated
NVIDIA RTX A600068.51 tok/s48GBMeasured
AMD Radeon Pro W790066.6 tok/s48GBEstimated
NVIDIA GeForce RTX 4070 Ti Super65.89 tok/s16GBMeasured
AMD Radeon RX 7900 XT65.1 tok/s20GBEstimated
NVIDIA RTX A500064.89 tok/s24GBMeasured
NVIDIA RTX PRO 4000 Blackwell60.62 tok/s24GBMeasured
NVIDIA Titan RTX59.75 tok/s24GBMeasured
AMD Radeon RX 9070 XT59.0 tok/s16GBEstimated
GeForce RTX 507058.15 tok/s12GBMeasured
NVIDIA A4057.85 tok/s48GBMeasured
NVIDIA RTX 5000 Ada Generation56.66 tok/s32GBMeasured
AMD Radeon RX 907056.6 tok/s16GBEstimated
NVIDIA TITAN V56.25 tok/s12GBMeasured
NVIDIA RTX A450054.29 tok/s20GBMeasured
AMD Radeon RX 7800 XT52.8 tok/s16GBEstimated
NVIDIA GeForce RTX 4070 Super51.23 tok/s12GBMeasured
NVIDIA GeForce RTX 407050.74 tok/s12GBMeasured
NVIDIA GeForce RTX 4070 Ti50.73 tok/s12GBMeasured
NVIDIA A10G47.07 tok/s24GBMeasured
NVIDIA Quadro RTX 800046.74 tok/s48GBMeasured
NVIDIA Quadro RTX 6000 (Turing)46.54 tok/s24GBMeasured
AMD Radeon RX 6900 XT45.8 tok/s16GBEstimated
GeForce RTX 5060 Ti45.06 tok/s16GBMeasured
AMD Radeon Pro W780044.1 tok/s32GBEstimated
AMD Radeon RX 6950 XT44.1 tok/s16GBEstimated
AMD Radeon Pro W680043.6 tok/s32GBEstimated
AMD Radeon RX 6800 XT43.6 tok/s16GBEstimated
AMD Radeon RX 680043.6 tok/s16GBEstimated
NVIDIA RTX 4500 Ada Generation43.38 tok/s24GBMeasured
NVIDIA RTX A400041.13 tok/s16GBMeasured
NVIDIA Quadro RTX 500041.11 tok/s16GBMeasured
AMD Radeon RX 7700 XT37.2 tok/s12GBEstimated
NVIDIA RTX 4000 (Ada Generation)36.34 tok/s20GBMeasured
NVIDIA TITAN Xp36.2 tok/s12GBEstimated
NVIDIA GeForce RTX 306035.52 tok/s12GBMeasured
Intel Arc B58033.0 tok/s12GBEstimated
NVIDIA GeForce RTX 4060 Ti 16GB31.12 tok/s16GBMeasured
NVIDIA TITAN X (Pascal)30.5 tok/s12GBEstimated
NVIDIA L427.46 tok/s24GBMeasured
GeForce GTX 1080 Ti27.36 tok/s11GBMeasured
Intel Arc A770 Limited Edition24.9 tok/s16GBEstimated
NVIDIA RTX 2000 Ada Generation23.35 tok/s16GBMeasured
NVIDIA T419.76 tok/s16GBMeasured
Intel Arc Pro A6018.5 tok/s12GBEstimated

Tap any column to sort. Measured = we rented and ran this card ourselves. Estimated = interpolated against our measured anchors, never blended silently.

Because this workload is bandwidth-bound, the ranking above tracks memory bandwidth far more closely than core counts or price. A card with fewer tensor cores and faster memory will beat a card with the opposite.

That's the reason we run twelve workloads instead of publishing one score. A GPU isn't fast or slow. It's fast at some things and gated out of others, and which of those matters depends entirely on what you're actually going to run.

Our verdict

NVIDIA B300 tops our Qwen2.5-Coder 14B leaderboard at 158.5 tok/s (measured), 757% of the way clear of the slowest card that still fits. But the number that decides most purchases isn't on the chart. It's the 24 cards that can't run Qwen2.5-Coder 14B at all. This is a bandwidth workload: buy memory speed, not tensor cores.

FAQ

What is the fastest GPU for Qwen2.5-Coder 14B?
NVIDIA B300, at 158.5 tok/s on our bench, a first-party measurement. It carries 288GB of VRAM. Of the 102 cards we have Qwen2.5-Coder 14B data for, 78 can run it at all.
How much VRAM do I need for Qwen2.5-Coder 14B?
~11.5GB at Q4_K_M, comfortable on a 12GB card, easy on 16GB.
Why does the Qwen2.5-Coder 14B ranking look different from your other benchmarks?
Because this workload is bandwidth-bound, the ranking above tracks memory bandwidth far more closely than core counts or price. A card with fewer tensor cores and faster memory will beat a card with the opposite. That's why we publish twelve separate workloads rather than one blended score, the ordering genuinely changes depending on the job.
Are these Qwen2.5-Coder 14B numbers measured or estimated?
Both, and every row says which. 62 of the 102 cards here were rented and run by us on the same harness. The remainder are anchored estimates interpolated per workload against those measurements. We never blend the two silently, if a row says Estimated, we have not run that card.
Can I rent a GPU to run Qwen2.5-Coder 14B instead of buying one?
Yes, and for the cards at the top of this leaderboard it's the only realistic option, most of them have no retail channel at all. It's also how we got these numbers: we rented the hardware by the hour rather than buying it. That's worth considering before you spend on a card to find out whether it's fast enough.
Why publish cards that can't run Qwen2.5-Coder 14B?
Because it's the most useful thing we know. A card that can't load a model doesn't run it slowly, it doesn't run it. Most benchmark sites leave that as a blank cell or quietly drop to a smaller quantisation to produce a number. We publish it as a hard gate and score it zero, because 'this card cannot do the thing you want' is the answer to the question you were actually asking.

How we test

Every ranking on this page comes from our own benchmark runs, not vendor claims. Cards marked Measured were rented and run by us; cards marked Estimated are interpolated per workload against those measured anchors and are labelled on every row, we never blend the two silently. LLMs run on llama.cpp (llama-bench) at Q4_K_M with -p 512 -n 128. Diffusion and video run on diffusers/ComfyUI at BF16, with SDXL at FP16. Each workload gets a warmup pass plus multiple timed runs (5 for small LLMs, 3 for large models and images, 2 for video); we publish the mean as the result and the minimum as the 1% low. Run-to-run variance is under 0.5%. Telemetry, power, temperature, utilisation, clocks, peak VRAM, is sampled at 1 Hz for the duration of every run. Where a model exceeds a card's VRAM we publish a hard won't-fit result rather than quietly dropping to a smaller quantisation. A card that can't run a model scores zero on it. Silently swapping precision to make a number appear would make every number on this site meaningless. All figures are single-GPU, single-stream, batch-size-1. That is the honest way to measure what one card does for one user, and it is deliberately not how a datacenter serves a model. Vendor and MLPerf figures use large batches across many GPUs and will be far higher. Neither is wrong, they answer different questions. Ours answers 'what will this card do for me'.