Can You Run Llama-3.2-3B-Instruct-uncensored? Every GPU, Tested

A model that doesn't fit doesn't run slowly, it doesn't run. Here's every card from our bench, sorted by measured speed.

Weights: bartowski/Llama-3.2-3B-Instruct-uncensored-GGUF

GPUs that run Llama-3.2-3B-Instruct-uncensored (11)

GPUSpeedVRAMBasis
NVIDIA B300448.94 tok/s288GB✓ Measured
NVIDIA RTX PRO 6000 Blackwell Workstation Edition440.15 tok/s96GB✓ Measured
NVIDIA H200429.78 tok/s141GB✓ Measured
NVIDIA H100 80GB HBM3424.25 tok/s80GB✓ Measured
NVIDIA B200419.24 tok/s192GB✓ Measured
NVIDIA L40S271.01 tok/s48GB✓ Measured
NVIDIA A100 80GB SXM4269.13 tok/s80GB✓ Measured
NVIDIA A100 40GB SXM4260.83 tok/s40GBEst.
NVIDIA A10G167.63 tok/s24GB✓ Measured
NVIDIA L4101.16 tok/s24GB✓ Measured
NVIDIA T486.29 tok/s16GB✓ Measured

Quick answers

What is the fastest GPU for Llama-3.2-3B-Instruct-uncensored?

NVIDIA B300 is the fastest card we've measured running Llama-3.2-3B-Instruct-uncensored, at 448.94 tok/s.

What is the cheapest GPU that can run Llama-3.2-3B-Instruct-uncensored?

NVIDIA T4 ($2,299 MSRP) is the cheapest card on our bench that runs Llama-3.2-3B-Instruct-uncensored, at 86.29 tok/s.

← All models · AI GPU rankings · How we benchmark