Llama-3.3-70B-Instruct-abliterated needs roughly 54GB of VRAM at our tested precision. A model that doesn't fit doesn't run slowly, it doesn't run. Here's every card from our bench, sorted by measured speed.
Weights: bartowski/Llama-3.3-70B-Instruct-abliterated-GGUF
| GPU | Speed | VRAM | Basis |
|---|---|---|---|
| NVIDIA B300 | 48.1 tok/s | 288GB | ✓ Measured |
| NVIDIA B200 | 44.48 tok/s | 192GB | ✓ Measured |
| NVIDIA H200 | 42.75 tok/s | 141GB | ✓ Measured |
| NVIDIA H100 80GB HBM3 | 41.15 tok/s | 80GB | ✓ Measured |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 32.53 tok/s | 96GB | ✓ Measured |
| NVIDIA A100 80GB SXM4 | 22.92 tok/s | 80GB | ✓ Measured |
| NVIDIA L40S | 16.49 tok/s | 48GB | ✓ Measured |
These cards' VRAM is below the requirement at tested precision, published as data, not hidden.
| GPU | VRAM |
|---|---|
| NVIDIA A100 40GB SXM4 | 40GB |
| NVIDIA A10G | 24GB |
| NVIDIA L4 | 24GB |
| NVIDIA T4 | 16GB |
Llama-3.3-70B-Instruct-abliterated needs roughly 54GB of VRAM at our tested precision. 7 of the GPUs on our bench run it; 4 don't have the VRAM for it.
NVIDIA B300 is the fastest card we've measured running Llama-3.3-70B-Instruct-abliterated, at 48.1 tok/s.
NVIDIA L40S ($7,500 MSRP) is the cheapest card on our bench that runs Llama-3.3-70B-Instruct-abliterated, at 16.49 tok/s.