Can You Run Qwen2.5-Coder 32B (Q3_K_M)? Every GPU, Tested

Qwen2.5-Coder 32B (Q3_K_M) needs roughly 16GB of VRAM at our tested precision. A model that doesn't fit doesn't run slowly, it doesn't run. Here's every card from our bench, sorted by measured speed.

Weights: bartowski/Qwen2.5-Coder-32B-Instruct-GGUF

GPUs that run Qwen2.5-Coder 32B (Q3_K_M) (10)

GPUSpeedVRAMBasis
NVIDIA RTX PRO 6000 Blackwell Workstation Edition75.78 tok/s96GB✓ Measured
NVIDIA B30074.28 tok/s288GB✓ Measured
NVIDIA B20060.24 tok/s192GB✓ Measured
NVIDIA H20058.87 tok/s141GB✓ Measured
NVIDIA H100 80GB HBM358.05 tok/s80GB✓ Measured
NVIDIA L40S41.06 tok/s48GB✓ Measured
NVIDIA A100 80GB SXM431.99 tok/s80GB✓ Measured
NVIDIA A100 40GB SXM431.04 tok/s40GBEst.
NVIDIA A10G18.78 tok/s24GB✓ Measured
NVIDIA L412.05 tok/s24GB✓ Measured

GPUs it won't fit on (1)

These cards' VRAM is below the requirement at tested precision, published as data, not hidden.

GPUVRAM
NVIDIA T416GB

Quick answers

How much VRAM does Qwen2.5-Coder 32B (Q3_K_M) need?

Qwen2.5-Coder 32B (Q3_K_M) needs roughly 16GB of VRAM at our tested precision. 10 of the GPUs on our bench run it; 1 don't have the VRAM for it.

What is the fastest GPU for Qwen2.5-Coder 32B (Q3_K_M)?

NVIDIA RTX PRO 6000 Blackwell Workstation Edition is the fastest card we've measured running Qwen2.5-Coder 32B (Q3_K_M), at 75.78 tok/s.

What is the cheapest GPU that can run Qwen2.5-Coder 32B (Q3_K_M)?

NVIDIA L4 ($2,500 MSRP) is the cheapest card on our bench that runs Qwen2.5-Coder 32B (Q3_K_M), at 12.05 tok/s.

← All models · AI GPU rankings · How we benchmark