Can You Run Llama-3.3-70B-Instruct-abliterated? Every GPU, Tested

Llama-3.3-70B-Instruct-abliterated needs roughly 54GB of VRAM at our tested precision. A model that doesn't fit doesn't run slowly, it doesn't run. Here's every card from our bench, sorted by measured speed.

Weights: bartowski/Llama-3.3-70B-Instruct-abliterated-GGUF

GPUs that run Llama-3.3-70B-Instruct-abliterated (7)

GPUSpeedVRAMBasis
NVIDIA B30048.1 tok/s288GB✓ Measured
NVIDIA B20044.48 tok/s192GB✓ Measured
NVIDIA H20042.75 tok/s141GB✓ Measured
NVIDIA H100 80GB HBM341.15 tok/s80GB✓ Measured
NVIDIA RTX PRO 6000 Blackwell Workstation Edition32.53 tok/s96GB✓ Measured
NVIDIA A100 80GB SXM422.92 tok/s80GB✓ Measured
NVIDIA L40S16.49 tok/s48GB✓ Measured

GPUs it won't fit on (4)

These cards' VRAM is below the requirement at tested precision, published as data, not hidden.

GPUVRAM
NVIDIA A100 40GB SXM440GB
NVIDIA A10G24GB
NVIDIA L424GB
NVIDIA T416GB

Quick answers

How much VRAM does Llama-3.3-70B-Instruct-abliterated need?

Llama-3.3-70B-Instruct-abliterated needs roughly 54GB of VRAM at our tested precision. 7 of the GPUs on our bench run it; 4 don't have the VRAM for it.

What is the fastest GPU for Llama-3.3-70B-Instruct-abliterated?

NVIDIA B300 is the fastest card we've measured running Llama-3.3-70B-Instruct-abliterated, at 48.1 tok/s.

What is the cheapest GPU that can run Llama-3.3-70B-Instruct-abliterated?

NVIDIA L40S ($7,500 MSRP) is the cheapest card on our bench that runs Llama-3.3-70B-Instruct-abliterated, at 16.49 tok/s.

← All models · AI GPU rankings · How we benchmark