Can You Run Qwen2.5-Coder-7B-Instruct-abliterated? Every GPU, Tested

A model that doesn't fit doesn't run slowly, it doesn't run. Here's every card from our bench, sorted by measured speed.

Weights: bartowski/Qwen2.5-Coder-7B-Instruct-abliterated-GGUF

GPUs that run Qwen2.5-Coder-7B-Instruct-abliterated (11)

GPUSpeedVRAMBasis
NVIDIA B300292.21 tok/s288GB✓ Measured
NVIDIA B200276.57 tok/s192GB✓ Measured
NVIDIA H200270.92 tok/s141GB✓ Measured
NVIDIA H100 80GB HBM3264.24 tok/s80GB✓ Measured
NVIDIA RTX PRO 6000 Blackwell Workstation Edition257.07 tok/s96GB✓ Measured
NVIDIA A100 80GB SXM4166.84 tok/s80GB✓ Measured
NVIDIA A100 40GB SXM4159.76 tok/s40GBEst.
NVIDIA L40S142.97 tok/s48GB✓ Measured
NVIDIA A10G91.41 tok/s24GB✓ Measured
NVIDIA L453.02 tok/s24GB✓ Measured
NVIDIA T437.31 tok/s16GB✓ Measured

Quick answers

What is the fastest GPU for Qwen2.5-Coder-7B-Instruct-abliterated?

NVIDIA B300 is the fastest card we've measured running Qwen2.5-Coder-7B-Instruct-abliterated, at 292.21 tok/s.

What is the cheapest GPU that can run Qwen2.5-Coder-7B-Instruct-abliterated?

NVIDIA T4 ($2,299 MSRP) is the cheapest card on our bench that runs Qwen2.5-Coder-7B-Instruct-abliterated, at 37.31 tok/s.

← All models · AI GPU rankings · How we benchmark