Dolphin 3.0 Llama 3.1 8B · 11 GPUs measured first-party · llama.cpp Q4_K_M · Updated October 2026

What GPU Do You Need for Dolphin 3.0 Llama 3.1 8B?

Dolphin 3.0 Llama 3.1 8B puts the uncensored Dolphin treatment on Meta's Llama 3.1 base, an everyday-size assistant that answers without alignment hedging. Measured on 11 GPUs (llama.cpp, Q4_K_M): 286 tok/s on the B300, ~6GB peak VRAM, 2.49 tok/W best-case efficiency.

Benchmarked weights: dphn/Dolphin3.0-Llama3.1-8B-GGUF

Fastest we measured
NVIDIA B300

NVIDIA B300

285.9 tok/s on Dolphin 3.0 Llama 3.1 8B, the ceiling. Measured on our bench. 288GB of VRAM, $40,000 at launch.

Pros
  • 285.9 tok/s on Dolphin 3.0 Llama 3.1 8B
  • 288GB, clears the Dolphin 3.0 Llama 3.1 8B floor
  • Rentable by the hour rather than bought
Cons
  • 1400W board rating
  • Datacenter or workstation hardware, not a retail purchase
Cheapest card that runs it
NVIDIA RTX PRO 6000 Blackwell Workstation Edition

NVIDIA RTX PRO 6000 Blackwell Workstation Edition

239.3 tok/s on Dolphin 3.0 Llama 3.1 8B, lowest launch price that still fits. Measured on our bench. 96GB of VRAM, $8,565 at launch.

Pros
  • 239.3 tok/s on Dolphin 3.0 Llama 3.1 8B
  • 96GB, clears the Dolphin 3.0 Llama 3.1 8B floor
  • Rentable by the hour rather than bought
Cons
  • 600W board rating
  • Datacenter or workstation hardware, not a retail purchase
285.9tok/s
Fastest: NVIDIA B300
measured
11
Cards that run Dolphin 3.0 Llama 3.1 8B
of 11 we have data for
0
Cards that can't run it at all
published as hard gates, not omissions
701%
Fastest vs slowest that fits
285.9 vs 35.68 tok/s

What GPU Do You Need for Dolphin 3.0 Llama 3.1 8B?, tok/s by GPU

NVIDIA B300
285.9 tok/s
NVIDIA B200
274.1 tok/s
NVIDIA H200
267.8 tok/s
NVIDIA H100 80GB HBM3
266.2 tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
239.3 tok/s
NVIDIA A100 80GB SXM4
158.6 tok/s
NVIDIA A100 40GB SXM4
156.6 tok/s
NVIDIA L40S
135.6 tok/s
NVIDIA A10G
86.54 tok/s
NVIDIA L4
50.24 tok/s
NVIDIA T4
35.68 tok/s

Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.

Efficiency: tok/s per 100W drawn

NVIDIA H100 80GB HBM3
248.54 tok/s / 100W
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
133.53 tok/s / 100W
NVIDIA H200
119.73 tok/s / 100W
NVIDIA A100 40GB SXM4
113.2 tok/s / 100W
NVIDIA B300
95.26 tok/s / 100W
NVIDIA A100 80GB SXM4
91.38 tok/s / 100W
NVIDIA L4
79.75 tok/s / 100W
NVIDIA B200
75.52 tok/s / 100W
NVIDIA A10G
71.11 tok/s / 100W
NVIDIA L40S
66.67 tok/s / 100W
NVIDIA T4
58.11 tok/s / 100W

Power is the average pulled during the run, sampled at 1Hz. The fastest card is often not the one here, and for anything left running this is the number that shows up on the bill.

Value: tok/s per $1,000 of MSRP

NVIDIA A10G
30.91 tok/s / $1k
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
27.94 tok/s / $1k
NVIDIA L4
20.1 tok/s / $1k
NVIDIA L40S
18.07 tok/s / $1k
NVIDIA T4
15.52 tok/s / $1k
NVIDIA A100 40GB SXM4
13.05 tok/s / $1k
NVIDIA A100 80GB SXM4
9.33 tok/s / $1k
NVIDIA H100 80GB HBM3
8.87 tok/s / $1k
NVIDIA H200
8.64 tok/s / $1k
NVIDIA B300
7.15 tok/s / $1k
NVIDIA B200
6.85 tok/s / $1k

Launch price, not street price, so it ages. A speed leaderboard always crowns the most expensive card; this is the counterweight.

Dolphin 3.0 Llama 3.1 8B. Measured generation speed by GPU

NVIDIA B300285.9
NVIDIA B200274.1
NVIDIA H200267.8
NVIDIA H100 80GB HBM3266.2
NVIDIA RTX PRO 6000 Blackwell Workstation Edition239.3
NVIDIA A100 80GB SXM4158.6
NVIDIA A100 40GB SXM4156.6
NVIDIA L40S135.6
NVIDIA A10G86.54
NVIDIA L450.24
NVIDIA T435.68
GPUtok/sPrompt t/stok/WAvg power
NVIDIA B300285.95286.30.95300.1 W
NVIDIA B200274.19689.60.76362.9 W
NVIDIA H200267.889491.2223.7 W
NVIDIA H100 80GB HBM3266.29047.72.49107.1 W
NVIDIA RTX PRO 6000 Blackwell Workstation Edition239.312768.21.34179.2 W
NVIDIA A100 80GB SXM4158.64687.50.91173.6 W
NVIDIA A100 40GB SXM4156.64317.31.13138.3 W
NVIDIA L40S135.69783.20.67203.3 W
NVIDIA A10G86.5431720.71121.7 W
NVIDIA L450.242946.80.863.0 W
NVIDIA T435.681174.40.5861.4 W

The daily-driver uncensored 8B. Where we slot Dolphin X1 as a specialist judge, this one is the generalist: a Llama 3.1 foundation, the most widely-tuned base in open source, with the censorship sanded off. It's the model to reach for when you want ordinary assistant work (drafting, summarizing, Q&A) from something that won't lecture you or quietly reshape its answers. And in a review role, same argument as the whole Dolphin line: judges should be blunt, and alignment training makes models diplomatic. Worth knowing before you commit: in our experience the newer Dolphin generations, this 3.0 included, are more neutral than the early releases were. The original Dolphins on older bases had a rawer edge that the modern tunes have smoothed out. Still clearly uncensored by aligned-model standards, but the gap has narrowed over time.

Bench behavior. 286 tok/s peak and ~6GB measured put it dead level with every other 8B we run, the uncensored tune is behaviorally different, not computationally different. The efficiency line worth noting is the L4: 50 tok/s at 48.8W, comfortably interactive for a single user on a card that sips power. This tier is the cheapest 'real assistant' hardware there is.

How it compares. H100 80GB HBM3: Dolphin 3.0 Llama 3.1 8B 266.2 tok/s, Dolphin X1 8B 266.0, Qwen2.5-VL 7B Instruct 267.7 (8B), DeepSeek-R1 Distill 7B 264.4, Llama 3 8B 264.4 (8B). 1 of 4 beat Dolphin 3.0 Llama 3.1 8B here.

Cost on a rented GPU. 1M generated tokens of Dolphin 3.0 Llama 3.1 8B: $0.84 on a A100 40GB SXM4 ($0.47/hr, 106 min), $6.74 on a B300 ($6.94/hr, 58 min, 8.1x the cost).

Dolphin 3.0 Llama 3.1 8B: cost per 1M generated tokens on rented GPUs

NVIDIA A100 40GB SXM4$0.47/hr
NVIDIA T4$0.14/hr
NVIDIA RTX PRO 6000 Blackwell Workstation Edition$1.08/hr
NVIDIA L40S$0.79/hr
NVIDIA A100 80GB SXM4$0.95/hr
NVIDIA H100 80GB HBM3$2.14/hr
NVIDIA L4$0.44/hr
NVIDIA H200$3.59/hr
NVIDIA B200$5.98/hr
NVIDIA B300$6.94/hr
GPUCheapest rateSpeed (tok/s)Cost per 1M generated tokens
NVIDIA A100 40GB SXM4$0.47/hr156.6$0.84
NVIDIA T4$0.14/hr35.68$1.06
NVIDIA RTX PRO 6000 Blackwell Workstation Edition$1.08/hr239.3$1.25
NVIDIA L40S$0.79/hr135.6$1.62
NVIDIA A100 80GB SXM4$0.95/hr158.6$1.66
NVIDIA H100 80GB HBM3$2.14/hr266.2$2.23
NVIDIA L4$0.44/hr50.24$2.43
NVIDIA H200$3.59/hr267.8$3.72
NVIDIA B200$5.98/hr274.1$6.06
NVIDIA B300$6.94/hr285.9$6.74

Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.

Speed tiers for Dolphin 3.0 Llama 3.1 8B. 30+ tok/s: 11 (B300, B200, H200). 30 tok/s is roughly where replies outpace reading.

Reading your prompt. Before Dolphin 3.0 Llama 3.1 8B writes anything it reads the input: 12768.2 tok/s on the RTX PRO 6000 Blackwell Workstation Edition (0.3s for a 4,000-token prompt), 1174.4 on the T4 (3.4s). Long documents and big code files feel this number more than the generation speed.

VRAM for Dolphin 3.0 Llama 3.1 8B. Measured peak 5.2GB, so 8GB is the smallest common card size; smallest card it ran on: T4 (16GB). With long context: Q4_K_M 6GB (tested), Q2_K 4GB, Q3_K_M 5GB, Q5_K_M 7GB, Q6_K 9GB.

Power on Dolphin 3.0 Llama 3.1 8B. Most efficient: RTX PRO 6000 Blackwell Workstation Edition, 179W, 0.21 kWh per 1M generated tokens. Hungriest: B200, 363W, 0.37 kWh. At $0.15/kWh: $0.031 per 1M generated tokens.

Our verdict

Dolphin 3.0 Llama 3.1 8B: 286 tok/s peak, ~6GB floor, a no-hedging generalist on the most familiar base in open source. Run it as your everyday uncensored assistant on any 8GB card, or as the blunt second reviewer in a consensus stack.

FAQ

What's the difference between this and stock Llama 3.1 8B?
The Dolphin fine-tune removes alignment-driven refusals and hedging. Hardware behavior is identical in our tests. The difference is that answers are direct and critiques uncushioned.
What GPU do I need?
8GB. Measured peak was ~6GB at Q4_K_M. From a $179 Arc A580 to a rented H200 (268 tok/s at 224W), the whole range runs it well.
Is it good as a consensus judge?
Yes. That's the Dolphin family's best trick. Uncensored models review other models' answers without diplomatic softening. This one's Llama base also adds lineage diversity if your panel is Qwen-heavy.
How fast is it on budget hardware?
50 tok/s on an L4 at under 50W, 36 tok/s on a T4, both comfortably faster than reading speed. For an 8B assistant, nearly any modern GPU is enough.
Dolphin 3.0 or the 24B Dolphins?
This for speed and cheap hardware; the Mistral-based 24Bs when you want the judge itself to be smarter (they need ~15GB but run near 121 tok/s peak). Same philosophy, bigger brain, one weight class up.
What quantization and settings did you test?
Q4_K_M via llama.cpp llama-bench, 5-run pp512+tg128 protocol with logged power and VRAM, the identical pinned harness as every LLM on this site, so its 286 tok/s peak compares one-to-one against the other 8Bs in our database.