Qwen2.5-Coder 32B (Q3_K_M) needs roughly 16GB of VRAM at our tested precision. A model that doesn't fit doesn't run slowly, it doesn't run. Here's every card from our bench, sorted by measured speed.
Weights: bartowski/Qwen2.5-Coder-32B-Instruct-GGUF
| GPU | Speed | VRAM | Basis |
|---|---|---|---|
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 75.78 tok/s | 96GB | ✓ Measured |
| NVIDIA B300 | 74.28 tok/s | 288GB | ✓ Measured |
| NVIDIA B200 | 60.24 tok/s | 192GB | ✓ Measured |
| NVIDIA H200 | 58.87 tok/s | 141GB | ✓ Measured |
| NVIDIA H100 80GB HBM3 | 58.05 tok/s | 80GB | ✓ Measured |
| NVIDIA L40S | 41.06 tok/s | 48GB | ✓ Measured |
| NVIDIA A100 80GB SXM4 | 31.99 tok/s | 80GB | ✓ Measured |
| NVIDIA A100 40GB SXM4 | 31.04 tok/s | 40GB | Est. |
| NVIDIA A10G | 18.78 tok/s | 24GB | ✓ Measured |
| NVIDIA L4 | 12.05 tok/s | 24GB | ✓ Measured |
These cards' VRAM is below the requirement at tested precision, published as data, not hidden.
| GPU | VRAM |
|---|---|
| NVIDIA T4 | 16GB |
Qwen2.5-Coder 32B (Q3_K_M) needs roughly 16GB of VRAM at our tested precision. 10 of the GPUs on our bench run it; 1 don't have the VRAM for it.
NVIDIA RTX PRO 6000 Blackwell Workstation Edition is the fastest card we've measured running Qwen2.5-Coder 32B (Q3_K_M), at 75.78 tok/s.
NVIDIA L4 ($2,500 MSRP) is the cheapest card on our bench that runs Qwen2.5-Coder 32B (Q3_K_M), at 12.05 tok/s.