Dolphin X1 8B · 11 GPUs measured first-party · llama.cpp Q4_K_M · Updated October 2026

What GPU Do You Need for Dolphin X1 8B?

Dolphin X1 8B is an uncensored fine-tune in the 8B class, and in our workflow it has a very specific job: the reviewer that says what it actually thinks. Measured on 11 GPUs (llama.cpp, Q4_K_M): 287 tok/s on the B300, ~6GB peak VRAM.

Benchmarked weights: dphn/Dolphin-X1-8B-GGUF

Fastest we measured
NVIDIA B300

NVIDIA B300

287.0 tok/s on Dolphin X1 8B, the ceiling. Measured on our bench. 288GB of VRAM, $40,000 at launch.

Pros
  • 287.0 tok/s on Dolphin X1 8B
  • 288GB, clears the Dolphin X1 8B floor
  • Rentable by the hour rather than bought
Cons
  • 1400W board rating
  • Datacenter or workstation hardware, not a retail purchase
Cheapest card that runs it
NVIDIA RTX PRO 6000 Blackwell Workstation Edition

NVIDIA RTX PRO 6000 Blackwell Workstation Edition

239.4 tok/s on Dolphin X1 8B, lowest launch price that still fits. Measured on our bench. 96GB of VRAM, $8,565 at launch.

Pros
  • 239.4 tok/s on Dolphin X1 8B
  • 96GB, clears the Dolphin X1 8B floor
  • Rentable by the hour rather than bought
Cons
  • 600W board rating
  • Datacenter or workstation hardware, not a retail purchase
287.0tok/s
Fastest: NVIDIA B300
measured
11
Cards that run Dolphin X1 8B
of 11 we have data for
0
Cards that can't run it at all
published as hard gates, not omissions
705%
Fastest vs slowest that fits
287.0 vs 35.66 tok/s

What GPU Do You Need for Dolphin X1 8B?, tok/s by GPU

NVIDIA B300
287 tok/s
NVIDIA B200
273.9 tok/s
NVIDIA H200
268.4 tok/s
NVIDIA H100 80GB HBM3
266 tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
239.4 tok/s
NVIDIA A100 80GB SXM4
158.6 tok/s
NVIDIA A100 40GB SXM4
156.3 tok/s
NVIDIA L40S
135.5 tok/s
NVIDIA A10G
86.61 tok/s
NVIDIA L4
50.12 tok/s
NVIDIA T4
35.66 tok/s

Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.

Efficiency: tok/s per 100W drawn

NVIDIA H100 80GB HBM3
145.28 tok/s / 100W
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
118.75 tok/s / 100W
NVIDIA H200
116.73 tok/s / 100W
NVIDIA A100 40GB SXM4
102.09 tok/s / 100W
NVIDIA A100 80GB SXM4
101.94 tok/s / 100W
NVIDIA B300
98.4 tok/s / 100W
NVIDIA B200
83.85 tok/s / 100W
NVIDIA L4
79.94 tok/s / 100W
NVIDIA A10G
71.7 tok/s / 100W
NVIDIA L40S
67.94 tok/s / 100W
NVIDIA T4
58.65 tok/s / 100W

Power is the average pulled during the run, sampled at 1Hz. The fastest card is often not the one here, and for anything left running this is the number that shows up on the bill.

Value: tok/s per $1,000 of MSRP

NVIDIA A10G
30.93 tok/s / $1k
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
27.95 tok/s / $1k
NVIDIA L4
20.05 tok/s / $1k
NVIDIA L40S
18.06 tok/s / $1k
NVIDIA T4
15.51 tok/s / $1k
NVIDIA A100 40GB SXM4
13.03 tok/s / $1k
NVIDIA A100 80GB SXM4
9.33 tok/s / $1k
NVIDIA H100 80GB HBM3
8.87 tok/s / $1k
NVIDIA H200
8.66 tok/s / $1k
NVIDIA B300
7.18 tok/s / $1k
NVIDIA B200
6.85 tok/s / $1k

Launch price, not street price, so it ages. A speed leaderboard always crowns the most expensive card; this is the counterweight.

Dolphin X1 8B. Measured generation speed by GPU

NVIDIA B300287
NVIDIA B200273.9
NVIDIA H200268.4
NVIDIA H100 80GB HBM3266
NVIDIA RTX PRO 6000 Blackwell Workstation Edition239.4
NVIDIA A100 80GB SXM4158.6
NVIDIA A100 40GB SXM4156.3
NVIDIA L40S135.5
NVIDIA A10G86.61
NVIDIA L450.12
NVIDIA T435.66
GPUtok/sPrompt t/stok/WAvg power
NVIDIA B3002875410.20.98291.7 W
NVIDIA B200273.99655.60.84326.6 W
NVIDIA H200268.49001.71.17229.9 W
NVIDIA H100 80GB HBM32668958.61.45183.1 W
NVIDIA RTX PRO 6000 Blackwell Workstation Edition239.4129921.19201.6 W
NVIDIA A100 80GB SXM4158.647051.02155.6 W
NVIDIA A100 40GB SXM4156.34327.21.02153.1 W
NVIDIA L40S135.59871.10.68199.4 W
NVIDIA A10G86.613171.40.72120.8 W
NVIDIA L450.122986.80.862.7 W
NVIDIA T435.661172.80.5960.8 W

Why uncensored matters for judging. Aligned models don't just refuse things: their safety training subtly shapes every output, softening critiques and hedging judgments. That's exactly the wrong property for a model whose job is to review another model's work. This is our case for keeping a Dolphin in the stack: when you run a consensus setup, make the judge uncensored, because you want the referee calling fouls, not being polite about them. Dolphin X1 8B is the lightest model we benchmark that fills that seat. One honest caveat from experience: the Dolphin line has drifted over the years. The early Dolphins built on the older open-source bases were genuinely blunt; the recent releases, X1 included, have grown noticeably more neutral. It's still far less filtered than an aligned model, just don't expect the old-school Dolphin edge.

What we measured. Speed-wise it's a standard 8B: 287 tok/s peak, the familiar tight top three, and a ~6GB floor that any 8GB card clears. The H100's 1.45 tok/W leads efficiency. There's no performance penalty for the uncensored tune, it benchmarks within noise of the other 8Bs in our database, so the decision to run it is purely about behavior, not hardware.

How it compares. H100 80GB HBM3: Dolphin X1 8B 266.0 tok/s, Dolphin 3.0 Llama 3.1 8B 266.2, DeepSeek-R1 Distill 7B 264.4, Llama 3 8B 264.4 (8B), Qwen2.5-VL 7B Instruct 267.7 (8B). 2 of 4 beat Dolphin X1 8B here.

Cost on a rented GPU. 1M generated tokens of Dolphin X1 8B: $0.84 on a A100 40GB SXM4 ($0.47/hr, 107 min), $6.72 on a B300 ($6.94/hr, 58 min, 8.0x the cost).

Dolphin X1 8B: cost per 1M generated tokens on rented GPUs

NVIDIA A100 40GB SXM4$0.47/hr
NVIDIA T4$0.14/hr
NVIDIA RTX PRO 6000 Blackwell Workstation Edition$1.08/hr
NVIDIA L40S$0.79/hr
NVIDIA A100 80GB SXM4$0.95/hr
NVIDIA H100 80GB HBM3$2.14/hr
NVIDIA L4$0.44/hr
NVIDIA H200$3.59/hr
NVIDIA B200$5.98/hr
NVIDIA B300$6.94/hr
GPUCheapest rateSpeed (tok/s)Cost per 1M generated tokens
NVIDIA A100 40GB SXM4$0.47/hr156.3$0.84
NVIDIA T4$0.14/hr35.66$1.06
NVIDIA RTX PRO 6000 Blackwell Workstation Edition$1.08/hr239.4$1.25
NVIDIA L40S$0.79/hr135.5$1.62
NVIDIA A100 80GB SXM4$0.95/hr158.6$1.66
NVIDIA H100 80GB HBM3$2.14/hr266$2.23
NVIDIA L4$0.44/hr50.12$2.44
NVIDIA H200$3.59/hr268.4$3.72
NVIDIA B200$5.98/hr273.9$6.07
NVIDIA B300$6.94/hr287$6.72

Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.

Speed tiers for Dolphin X1 8B. 30+ tok/s: 11 (B300, B200, H200). 30 tok/s is roughly where replies outpace reading.

Reading your prompt. Before Dolphin X1 8B writes anything it reads the input: 12992.0 tok/s on the RTX PRO 6000 Blackwell Workstation Edition (0.3s for a 4,000-token prompt), 1172.8 on the T4 (3.4s). Long documents and big code files feel this number more than the generation speed.

VRAM for Dolphin X1 8B. Measured peak 5.2GB, so 8GB is the smallest common card size; smallest card it ran on: T4 (16GB). With long context: Q4_K_M 6GB (tested), Q5_K_M 7GB, Q6_K 9GB, Q8_0 11GB.

Power on Dolphin X1 8B. Most efficient: H100 80GB HBM3, 183W, 0.19 kWh per 1M generated tokens. Hungriest: B200, 327W, 0.33 kWh. At $0.15/kWh: $0.029 per 1M generated tokens.

Our verdict

Dolphin X1 8B: 287 tok/s peak, ~6GB floor, hardware-identical to every other 8B we've measured. Its value is behavioral: an uncensored reviewer whose critiques aren't softened by alignment training, the referee seat in a consensus stack, on any 8GB card.

FAQ

What is Dolphin X1 8B actually for?
Reviewing and judging. Uncensored tunes give direct, unhedged critiques, which makes them the right referee when you run multiple models on one prompt and need an honest verdict on the answers.
What GPU does it need?
Any 8GB card, ~6GB measured peak at Q4_K_M. It benchmarks identically to other 8Bs: 287 tok/s on the B300 down to 34 tok/s on a T4.
Does the uncensored tune cost performance?
Not in our measurements, throughput and VRAM land within noise of Llama- and Qwen-based 8Bs. The difference is entirely in output behavior, not speed.
Is an uncensored model safe to build with?
It follows your instructions rather than a safety policy, so the guardrails become your responsibility. In a judge role, critiquing outputs privately in your pipeline, that tradeoff is usually easy to accept.
Dolphin X1 8B or Dolphin 3.0 Llama 3.1 8B?
Same size, same measured speed within a tokens-per-second of each other. Different bases and tuning vintages, so they critique with different flavors, try both as judge, keep the one whose reviews you trust.
How do I set up a consensus pipeline with it as judge?
Simplest working recipe: send the prompt to two answer models (say Qwen3 8B and DeepSeek-R1 Distill 7B), then hand both answers to Dolphin X1 with the instruction to compare, find faults, and pick or merge. All three are ~6GB models, so one 8GB card runs the whole panel sequentially in under a minute per question.