Can You Run DeepSeek-R1-Distill-Qwen-32B-abliterated? Every GPU, Tested

DeepSeek-R1-Distill-Qwen-32B-abliterated needs roughly 25GB of VRAM at our tested precision. A model that doesn't fit doesn't run slowly, it doesn't run. Here's every card from our bench, sorted by measured speed.

Weights: bartowski/DeepSeek-R1-Distill-Qwen-32B-abliterated-GGUF

GPUs that run DeepSeek-R1-Distill-Qwen-32B-abliterated (10)

GPUSpeedVRAMBasis
NVIDIA B30083.19 tok/s288GB✓ Measured
NVIDIA B20078.48 tok/s192GB✓ Measured
NVIDIA H20075.71 tok/s141GB✓ Measured
NVIDIA H100 80GB HBM373.48 tok/s80GB✓ Measured
NVIDIA RTX PRO 6000 Blackwell Workstation Edition66.03 tok/s96GB✓ Measured
NVIDIA A100 80GB SXM443.48 tok/s80GB✓ Measured
NVIDIA A100 40GB SXM443.07 tok/s40GBEst.
NVIDIA L40S34.43 tok/s48GB✓ Measured
NVIDIA A10G22.14 tok/s24GB✓ Measured
NVIDIA L412.26 tok/s24GB✓ Measured

GPUs it won't fit on (1)

These cards' VRAM is below the requirement at tested precision, published as data, not hidden.

GPUVRAM
NVIDIA T416GB

Quick answers

How much VRAM does DeepSeek-R1-Distill-Qwen-32B-abliterated need?

DeepSeek-R1-Distill-Qwen-32B-abliterated needs roughly 25GB of VRAM at our tested precision. 10 of the GPUs on our bench run it; 1 don't have the VRAM for it.

What is the fastest GPU for DeepSeek-R1-Distill-Qwen-32B-abliterated?

NVIDIA B300 is the fastest card we've measured running DeepSeek-R1-Distill-Qwen-32B-abliterated, at 83.19 tok/s.

What is the cheapest GPU that can run DeepSeek-R1-Distill-Qwen-32B-abliterated?

NVIDIA L4 ($2,500 MSRP) is the cheapest card on our bench that runs DeepSeek-R1-Distill-Qwen-32B-abliterated, at 12.26 tok/s.

← All models · AI GPU rankings · How we benchmark