Qwen2.5-Coder 7B · 24 GPUs measured first-party · llama.cpp Q4_K_M · Updated October 2026

What GPU Do You Need for Qwen2.5-Coder 7B?

Qwen2.5-Coder 7B is the autocomplete tier: small enough to fit an 8GB card (~6GB measured peak), fast enough that completions feel instant: 287 tok/s on the B300, 266 on the H200, and still 38 tok/s on a seven-year-old T4. We measured it on 24 GPUs with llama.cpp at Q4_K_M, logging power and VRAM on every run.

Benchmarked weights: bartowski/Qwen2.5-Coder-7B-Instruct-GGUF

Fastest we measured
NVIDIA B300

NVIDIA B300

286.5 tok/s on Qwen2.5-Coder 7B, the ceiling. Measured on our bench. 288GB of VRAM, $40,000 at launch.

Pros
  • 286.5 tok/s on Qwen2.5-Coder 7B
  • 288GB, clears the Qwen2.5-Coder 7B floor
  • Rentable by the hour rather than bought
Cons
  • 1400W board rating
  • Datacenter or workstation hardware, not a retail purchase
Best consumer card
NVIDIA GeForce RTX 5090

NVIDIA GeForce RTX 5090

284.6 tok/s on Qwen2.5-Coder 7B, fastest card you can buy at retail. Measured on our bench. 32GB of VRAM, $1,999 at launch.

Pros
  • 284.6 tok/s on Qwen2.5-Coder 7B
  • 32GB, clears the Qwen2.5-Coder 7B floor
Cons
  • 575W board rating
Cheapest card that runs it
NVIDIA GeForce GTX 1660 Super

NVIDIA GeForce GTX 1660 Super

49.04 tok/s on Qwen2.5-Coder 7B, lowest launch price that still fits. Measured on our bench. 6GB of VRAM, $229 at launch.

Pros
  • 49.04 tok/s on Qwen2.5-Coder 7B
  • 6GB, clears the Qwen2.5-Coder 7B floor
Cons
  • 125W board rating
Best value
NVIDIA GeForce RTX 5060

NVIDIA GeForce RTX 5060

81.99 tok/s on Qwen2.5-Coder 7B, most speed per dollar. Measured on our bench. 8GB of VRAM, $249 at launch. That is 329.3 tok/s per $1,000 of launch price.

Pros
  • 81.99 tok/s on Qwen2.5-Coder 7B
  • 8GB, clears the Qwen2.5-Coder 7B floor
Cons
  • 145W board rating
286.5tok/s
Fastest: NVIDIA B300
measured
11
Cards that run Qwen2.5-Coder 7B
of 11 we have data for
0
Cards that can't run it at all
published as hard gates, not omissions
663%
Fastest vs slowest that fits
286.5 vs 37.54 tok/s

What GPU Do You Need for Qwen2.5-Coder 7B?, tok/s by GPU

NVIDIA B300
286.5 tok/s
NVIDIA GeForce RTX 5090
284.6 tok/s
NVIDIA B200
278 tok/s
NVIDIA H200
266.4 tok/s
NVIDIA H100 80GB HBM3
263.9 tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
256.1 tok/s
NVIDIA GeForce RTX 4090
183.7 tok/s
GeForce RTX 5080
173.5 tok/s
NVIDIA A100 80GB SXM4
164.6 tok/s
GeForce RTX 5070 Ti
161 tok/s
NVIDIA A100 40GB SXM4
159.9 tok/s
NVIDIA GeForce RTX 3090
156.5 tok/s
NVIDIA L40S
143.9 tok/s
NVIDIA GeForce RTX 4080
134.4 tok/s
NVIDIA A10G
90.05 tok/s

Top 15 shown; 9 more cards in the full table below.

Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.

Efficiency: tok/s per 100W drawn

NVIDIA H200
202.09 tok/s / 100W
NVIDIA A100 80GB SXM4
144.64 tok/s / 100W
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
131.68 tok/s / 100W
NVIDIA H100 80GB HBM3
109.46 tok/s / 100W
NVIDIA GeForce RTX 5090
107.31 tok/s / 100W
GeForce RTX 5070 Ti
99.39 tok/s / 100W
NVIDIA B300
95.95 tok/s / 100W
GeForce RTX 5080
91.38 tok/s / 100W
NVIDIA A100 40GB SXM4
90.85 tok/s / 100W
NVIDIA L4
85.22 tok/s / 100W
NVIDIA GeForce RTX 4090
83.25 tok/s / 100W
NVIDIA B200
81.57 tok/s / 100W
NVIDIA GeForce RTX 5060
76.13 tok/s / 100W
GeForce RTX 5060 Ti
75.42 tok/s / 100W
NVIDIA A10G
74.24 tok/s / 100W

Top 15 shown; 9 more cards in the full table below.

Power is the average pulled during the run, sampled at 1Hz. The fastest card is often not the one here, and for anything left running this is the number that shows up on the bill.

Value: tok/s per $1,000 of MSRP

NVIDIA GeForce RTX 5060
329.28 tok/s / $1k
GeForce RTX 5070 Ti
214.97 tok/s / $1k
NVIDIA GeForce GTX 1660 Super
214.15 tok/s / $1k
NVIDIA GeForce RTX 3060
205.41 tok/s / $1k
GeForce RTX 5060 Ti
203.24 tok/s / $1k
NVIDIA GeForce RTX 2060 Super
179.5 tok/s / $1k
GeForce RTX 5080
173.7 tok/s / $1k
NVIDIA GeForce RTX 2070 SUPER
161.52 tok/s / $1k
NVIDIA GeForce RTX 5090
142.37 tok/s / $1k
NVIDIA GeForce RTX 4060 Ti 16GB
117.19 tok/s / $1k
NVIDIA GeForce RTX 4090
114.9 tok/s / $1k
NVIDIA GeForce RTX 4080
112.06 tok/s / $1k
NVIDIA GeForce RTX 3090
104.42 tok/s / $1k
NVIDIA A10G
32.16 tok/s / $1k
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
29.9 tok/s / $1k

Top 15 shown; 9 more cards in the full table below.

Launch price, not street price, so it ages. A speed leaderboard always crowns the most expensive card; this is the counterweight.

Qwen2.5-Coder 7B. Measured generation speed by GPU

NVIDIA B300286.5
NVIDIA GeForce RTX 5090284.6
NVIDIA B200278
NVIDIA H200266.4
NVIDIA H100 80GB HBM3263.9
NVIDIA RTX PRO 6000 Blackwell Workstation Edition256.1
NVIDIA GeForce RTX 4090183.7
GeForce RTX 5080173.5
NVIDIA A100 80GB SXM4164.6
GeForce RTX 5070 Ti161
NVIDIA A100 40GB SXM4159.9
NVIDIA GeForce RTX 3090156.5
NVIDIA L40S143.9
NVIDIA GeForce RTX 4080134.4
NVIDIA A10G90.05
GeForce RTX 5060 Ti87.19
NVIDIA GeForce RTX 506081.99
NVIDIA GeForce RTX 2070 SUPER80.6
NVIDIA GeForce RTX 2060 Super71.62
NVIDIA GeForce RTX 306067.58
NVIDIA GeForce RTX 4060 Ti 16GB58.48
NVIDIA L453.26
NVIDIA GeForce GTX 1660 Super49.04
NVIDIA T437.54
GPUtok/sPrompt t/stok/WAvg power
NVIDIA B300286.55832.30.96298.6 W
NVIDIA GeForce RTX 5090284.615688.81.07265.2 W
NVIDIA B20027810233.50.82340.8 W
NVIDIA H200266.48971.22.02131.8 W
NVIDIA H100 80GB HBM3263.99408.21.09241.1 W
NVIDIA RTX PRO 6000 Blackwell Workstation Edition256.113046.71.32194.5 W
NVIDIA GeForce RTX 4090183.712068.80.83220.7 W
GeForce RTX 5080173.58563.10.91189.9 W
NVIDIA A100 80GB SXM4164.64806.31.45113.8 W
GeForce RTX 5070 Ti1617400.40.99162.0 W
NVIDIA A100 40GB SXM4159.94736.10.91176.0 W
NVIDIA GeForce RTX 3090156.55975.80.53294.2 W
NVIDIA L40S143.910079.20.72199.5 W
NVIDIA GeForce RTX 4080134.48214.70.73183.2 W
NVIDIA A10G90.053546.60.74121.3 W
GeForce RTX 5060 Ti87.193904.40.75115.6 W
NVIDIA GeForce RTX 506081.9933070.76107.7 W
NVIDIA GeForce RTX 2070 SUPER80.62159.40.48168.5 W
NVIDIA GeForce RTX 2060 Super71.621704.10.48149.2 W
NVIDIA GeForce RTX 306067.582340.80.51132.8 W
NVIDIA GeForce RTX 4060 Ti 16GB58.483560.10.55107.2 W
NVIDIA L453.263194.70.8562.5 W
NVIDIA GeForce GTX 1660 Super49.04161.10.681.7 W
NVIDIA T437.541328.70.662.5 W

The right tool for inline completion. Code completion has a latency budget of a few hundred milliseconds. The model has to produce a usable suggestion before you type the next character. That's a speed problem, not an intelligence problem, and a 7B coder model is the honest answer to it. My recommended split for a local stack: this model handles fill-in-middle and autocomplete, while a 30B-class model (ideally the Qwen3 Coder MoE) handles the chat-and-refactor side. On a 24GB card you can run both simultaneously.

Numbers worth noticing. The H200's 266 tok/s at 132W (2.02 tok/W) doubles the efficiency of the chart-topping B300, the recurring Hopper-vs-Blackwell pattern in our small-model data. But the more practical row is further down: 6GB-class cards clear the floor, and the $179 Arc A580 fits it. For a completions model that runs all day next to your IDE, low idle-adjacent power draw matters more than peak throughput. This is a workload where a small efficient card is genuinely the better engineering choice.

About Qwen2.5-Coder 7B. Qwen2.5-Coder 7B: from Qwen, 7.6B parameters, on Hugging Face since September 2024, Apache 2.0 licence. 5,580,328 downloads in the last 30 days and 7 community quantizations.

How it compares. H100 80GB HBM3: Qwen2.5-Coder 7B 263.9 tok/s, Qwen2.5-7B 264.2 (8B), Mistral-7B-Instruct-v0.2 276.7 (7B), Llama-3.1-8B 261.8 (8B), Llama 3 8B 264.4 (8B). 3 of 4 beat Qwen2.5-Coder 7B here.

Cost on a rented GPU. 1M generated tokens of Qwen2.5-Coder 7B: $0.15 on a RTX 3060 ($0.036/hr, 4.1 hours), $6.73 on a B300 ($6.94/hr, 58 min, 45.5x the cost).

Qwen2.5-Coder 7B: cost per 1M generated tokens on rented GPUs

NVIDIA GeForce RTX 3060$0.036/hr
NVIDIA GeForce RTX 3090$0.12/hr
GeForce RTX 5070 Ti$0.15/hr
NVIDIA GeForce RTX 5060$0.090/hr
GeForce RTX 5080$0.21/hr
NVIDIA GeForce RTX 5090$0.39/hr
NVIDIA GeForce RTX 4080$0.20/hr
GeForce RTX 5060 Ti$0.14/hr
NVIDIA GeForce RTX 4090$0.34/hr
NVIDIA A100 40GB SXM4$0.47/hr
NVIDIA T4$0.14/hr
NVIDIA RTX PRO 6000 Blackwell Workstation Edition$1.08/hr
NVIDIA L40S$0.79/hr
NVIDIA A100 80GB SXM4$0.95/hr
NVIDIA H100 80GB HBM3$2.14/hr
NVIDIA L4$0.44/hr
NVIDIA H200$3.59/hr
NVIDIA B200$5.98/hr
NVIDIA B300$6.94/hr
GPUCheapest rateSpeed (tok/s)Cost per 1M generated tokens
NVIDIA GeForce RTX 3060$0.036/hr67.58$0.15
NVIDIA GeForce RTX 3090$0.12/hr156.5$0.22
GeForce RTX 5070 Ti$0.15/hr161$0.26
NVIDIA GeForce RTX 5060$0.090/hr81.99$0.30
GeForce RTX 5080$0.21/hr173.5$0.34
NVIDIA GeForce RTX 5090$0.39/hr284.6$0.38
NVIDIA GeForce RTX 4080$0.20/hr134.4$0.42
GeForce RTX 5060 Ti$0.14/hr87.19$0.43
NVIDIA GeForce RTX 4090$0.34/hr183.7$0.51
NVIDIA A100 40GB SXM4$0.47/hr159.9$0.82
NVIDIA T4$0.14/hr37.54$1.01
NVIDIA RTX PRO 6000 Blackwell Workstation Edition$1.08/hr256.1$1.17
NVIDIA L40S$0.79/hr143.9$1.52
NVIDIA A100 80GB SXM4$0.95/hr164.6$1.60
NVIDIA H100 80GB HBM3$2.14/hr263.9$2.25
NVIDIA L4$0.44/hr53.26$2.29
NVIDIA H200$3.59/hr266.4$3.74
NVIDIA B200$5.98/hr278$5.98
NVIDIA B300$6.94/hr286.5$6.73

Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.

Speed tiers for Qwen2.5-Coder 7B. 30+ tok/s: 24 (RTX 5090, RTX 4090, RTX 5080). 30 tok/s is roughly where replies outpace reading.

Reading your prompt. Before Qwen2.5-Coder 7B writes anything it reads the input: 15688.8 tok/s on the RTX 5090 (0.3s for a 4,000-token prompt), 12068.8 on the RTX 4090 (0.3s), 161.1 on the GTX 1660 Super (24.8s). Long documents and big code files feel this number more than the generation speed.

VRAM for Qwen2.5-Coder 7B. Measured peak 4.4GB, so 8GB is the smallest common card size; smallest card it ran on: GTX 1660 Super (6GB). With long context: Q4_K_M 6GB (tested), Q2_K 4GB, Q3_K_M 5GB, Q5_K_M 7GB, Q6_K 9GB.

Power on Qwen2.5-Coder 7B. Most efficient: A100 80GB SXM4, 114W, 0.19 kWh per 1M generated tokens. Hungriest: B200, 341W, 0.34 kWh. At $0.15/kWh: $0.029 per 1M generated tokens.

Our verdict

Qwen2.5-Coder 7B: 287 tok/s peak, ~6GB floor, instant-feel completions on almost any modern card. Use it as the fast half of a local coding stack, autocomplete here, a 30B-class model for the heavy lifting, and an 8GB card is all the hardware this half needs.

FAQ

Can an 8GB card run Qwen2.5-Coder 7B?
Comfortably, ~6GB measured peak at Q4_K_M leaves context headroom on 8GB. The $179 Intel Arc A580 clears it; so does an RTX 3050.
Is a 7B model good enough for coding?
For inline autocomplete and fill-in-middle, yes, that workload rewards speed over depth. For multi-file refactors and architectural questions, pair it with Qwen3 Coder 30B-A3B or another 30B-class model.
How fast is it in practice?
287 tok/s on the B300, 266 on the H200, but even the old T4's 38 tok/s outruns human reading speed. On any modern consumer card, completions are effectively instantaneous.
Qwen2.5-Coder 7B or the newer Qwen3 8B?
For code, the coder tune wins, fill-in-middle training matters for completions. Qwen3 8B is the better general assistant. Same VRAM class (~6GB), so your card runs either.
What does a full local coding stack look like?
Our suggested pairing: this model for completions plus Qwen3 Coder 30B-A3B for chat/agent work, together they fit in 24GB. On an 8GB card, run the 7B alone and let a cloud model handle the big tasks.