Dolphin 3.0 Llama 3.1 8B · 11 GPUs measured first-party · llama.cpp Q4_K_M · Updated July 2026

What GPU Do You Need for Dolphin 3.0 Llama 3.1 8B?

Dolphin 3.0 Llama 3.1 8B puts the uncensored Dolphin treatment on Meta's Llama 3.1 base, an everyday-size assistant that answers without alignment hedging. Measured on 11 GPUs (llama.cpp, Q4_K_M): 286 tok/s on the B300, ~6GB peak VRAM, 2.49 tok/W best-case efficiency.

Benchmarked weights: dphn/Dolphin3.0-Llama3.1-8B-GGUF

285.88tok/s
Fastest: NVIDIA B300
measured, 3-run llama-bench
~6GB
VRAM needed (measured peak)
GPU-independent, applies to every card
11
GPUs measured
same pinned harness
2.49tok/W
Most efficient: NVIDIA H100 80GB HBM3
real power sampling, not TDP

What GPU Do You Need for Dolphin 3.0 Llama 3.1 8B?, tok/s, fastest 11

NVIDIA B300
285.88 tok/s
NVIDIA B200
274.05 tok/s
NVIDIA H200
267.84 tok/s
NVIDIA H100 80GB HBM3
266.19 tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
239.28 tok/s
NVIDIA A100 80GB SXM4
158.63 tok/s
NVIDIA A100 40GB SXM4
157.99 tok/s
NVIDIA L40S
135.75 tok/s
NVIDIA A10G
87.9 tok/s
NVIDIA L4
50.26 tok/s
NVIDIA T4
35.55 tok/s

Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.

Dolphin 3.0 Llama 3.1 8B. Measured generation speed by GPU

NVIDIA B300285.88
NVIDIA B200274.05
NVIDIA H200267.84
NVIDIA H100 80GB HBM3266.19
NVIDIA RTX PRO 6000 Blackwell Workstation Edition239.28
NVIDIA A100 80GB SXM4158.63
NVIDIA A100 40GB SXM4157.99
NVIDIA L40S135.75
NVIDIA A10G87.9
NVIDIA L450.26
NVIDIA T435.55
GPUtok/sPrompt t/stok/WAvg power
NVIDIA B300285.885286.30.95300.1 W
NVIDIA B200274.059689.60.76362.9 W
NVIDIA H200267.8489491.2223.7 W
NVIDIA H100 80GB HBM3266.199047.72.49107.1 W
NVIDIA RTX PRO 6000 Blackwell Workstation Edition239.2812768.21.34179.2 W
NVIDIA A100 80GB SXM4158.634687.50.91173.6 W
NVIDIA A100 40GB SXM4157.9943171.08145.8 W
NVIDIA L40S135.759876.30.8168.9 W
NVIDIA A10G87.93263.9188.0 W
NVIDIA L450.262940.51.0348.8 W
NVIDIA T435.551182.10.7845.5 W

The daily-driver uncensored 8B. Where we slot Dolphin X1 as a specialist judge, this one is the generalist: a Llama 3.1 foundation, the most widely-tuned base in open source, with the censorship sanded off. It's the model to reach for when you want ordinary assistant work (drafting, summarizing, Q&A) from something that won't lecture you or quietly reshape its answers. And in a review role, same argument as the whole Dolphin line: judges should be blunt, and alignment training makes models diplomatic. Worth knowing before you commit: in our experience the newer Dolphin generations, this 3.0 included, are more neutral than the early releases were. The original Dolphins on older bases had a rawer edge that the modern tunes have smoothed out. Still clearly uncensored by aligned-model standards, but the gap has narrowed over time.

Bench behavior. 286 tok/s peak and ~6GB measured put it dead level with every other 8B we run, the uncensored tune is behaviorally different, not computationally different. The efficiency line worth noting is the L4: 50 tok/s at 48.8W, comfortably interactive for a single user on a card that sips power. This tier is the cheapest 'real assistant' hardware there is.

Our verdict

Dolphin 3.0 Llama 3.1 8B: 286 tok/s peak, ~6GB floor, a no-hedging generalist on the most familiar base in open source. Run it as your everyday uncensored assistant on any 8GB card, or as the blunt second reviewer in a consensus stack.

FAQ

What's the difference between this and stock Llama 3.1 8B?
The Dolphin fine-tune removes alignment-driven refusals and hedging. Hardware behavior is identical in our tests. The difference is that answers are direct and critiques uncushioned.
What GPU do I need?
8GB. Measured peak was ~6GB at Q4_K_M. From a $179 Arc A580 to a rented H200 (268 tok/s at 224W), the whole range runs it well.
Is it good as a consensus judge?
Yes. That's the Dolphin family's best trick. Uncensored models review other models' answers without diplomatic softening. This one's Llama base also adds lineage diversity if your panel is Qwen-heavy.
How fast is it on budget hardware?
50 tok/s on an L4 at under 50W, 36 tok/s on a T4, both comfortably faster than reading speed. For an 8B assistant, nearly any modern GPU is enough.
Dolphin 3.0 or the 24B Dolphins?
This for speed and cheap hardware; the Mistral-based 24Bs when you want the judge itself to be smarter (they need ~15GB but run near 121 tok/s peak). Same philosophy, bigger brain, one weight class up.
What quantization and settings did you test?
Q4_K_M via llama.cpp llama-bench, 5-run pp512+tg128 protocol with logged power and VRAM, the identical pinned harness as every LLM on this site, so its 286 tok/s peak compares one-to-one against the other 8Bs in our database.