Dolphin 3.0 Llama 3.1 8B · 11 GPUs measured first-party · llama.cpp Q4_K_M · Updated July 2026
Dolphin 3.0 Llama 3.1 8B puts the uncensored Dolphin treatment on Meta's Llama 3.1 base, an everyday-size assistant that answers without alignment hedging. Measured on 11 GPUs (llama.cpp, Q4_K_M): 286 tok/s on the B300, ~6GB peak VRAM, 2.49 tok/W best-case efficiency.
Benchmarked weights: dphn/Dolphin3.0-Llama3.1-8B-GGUF
What GPU Do You Need for Dolphin 3.0 Llama 3.1 8B?, tok/s, fastest 11
Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.
Dolphin 3.0 Llama 3.1 8B. Measured generation speed by GPU
| GPU | tok/s | Prompt t/s | tok/W | Avg power |
|---|---|---|---|---|
| NVIDIA B300 | 285.88 | 5286.3 | 0.95 | 300.1 W |
| NVIDIA B200 | 274.05 | 9689.6 | 0.76 | 362.9 W |
| NVIDIA H200 | 267.84 | 8949 | 1.2 | 223.7 W |
| NVIDIA H100 80GB HBM3 | 266.19 | 9047.7 | 2.49 | 107.1 W |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 239.28 | 12768.2 | 1.34 | 179.2 W |
| NVIDIA A100 80GB SXM4 | 158.63 | 4687.5 | 0.91 | 173.6 W |
| NVIDIA A100 40GB SXM4 | 157.99 | 4317 | 1.08 | 145.8 W |
| NVIDIA L40S | 135.75 | 9876.3 | 0.8 | 168.9 W |
| NVIDIA A10G | 87.9 | 3263.9 | 1 | 88.0 W |
| NVIDIA L4 | 50.26 | 2940.5 | 1.03 | 48.8 W |
| NVIDIA T4 | 35.55 | 1182.1 | 0.78 | 45.5 W |
The daily-driver uncensored 8B. Where we slot Dolphin X1 as a specialist judge, this one is the generalist: a Llama 3.1 foundation, the most widely-tuned base in open source, with the censorship sanded off. It's the model to reach for when you want ordinary assistant work (drafting, summarizing, Q&A) from something that won't lecture you or quietly reshape its answers. And in a review role, same argument as the whole Dolphin line: judges should be blunt, and alignment training makes models diplomatic. Worth knowing before you commit: in our experience the newer Dolphin generations, this 3.0 included, are more neutral than the early releases were. The original Dolphins on older bases had a rawer edge that the modern tunes have smoothed out. Still clearly uncensored by aligned-model standards, but the gap has narrowed over time.
Bench behavior. 286 tok/s peak and ~6GB measured put it dead level with every other 8B we run, the uncensored tune is behaviorally different, not computationally different. The efficiency line worth noting is the L4: 50 tok/s at 48.8W, comfortably interactive for a single user on a card that sips power. This tier is the cheapest 'real assistant' hardware there is.
Dolphin 3.0 Llama 3.1 8B: 286 tok/s peak, ~6GB floor, a no-hedging generalist on the most familiar base in open source. Run it as your everyday uncensored assistant on any 8GB card, or as the blunt second reviewer in a consensus stack.