Dolphin 3.0 R1 Mistral 24B · 10 GPUs measured first-party · llama.cpp Q4_K_M · Updated July 2026
Dolphin 3.0 R1 Mistral 24B is the most interesting model in the Dolphin line: an uncensored tune with R1-style reasoning on a Mistral 24B base, a judge that thinks before it criticizes. Measured on 10 GPUs (llama.cpp, Q4_K_M): 121 tok/s on the B300, ~15GB peak VRAM.
Benchmarked weights: bartowski/cognitivecomputations_Dolphin3.0-R1-Mistral-24B-GGUF
What GPU Do You Need for Dolphin 3.0 R1 Mistral 24B?, tok/s, fastest 10
Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.
Dolphin 3.0 R1 Mistral 24B. Measured generation speed by GPU
| GPU | tok/s | Prompt t/s | tok/W | Avg power |
|---|---|---|---|---|
| NVIDIA B300 | 120.95 | 1944 | 0.35 | 347.8 W |
| NVIDIA B200 | 113.95 | 3828 | 0.34 | 332.9 W |
| NVIDIA H200 | 109.06 | 3429.3 | 0.42 | 256.9 W |
| NVIDIA H100 80GB HBM3 | 107.83 | 3454.2 | 0.87 | 124.2 W |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 93.62 | 4902 | 0.47 | 199.1 W |
| NVIDIA A100 40GB SXM4 | 62.84 | 1642.6 | 0.47 | 132.8 W |
| NVIDIA A100 80GB SXM4 | 61.17 | 1673.7 | 0.33 | 186.6 W |
| NVIDIA L40S | 48.45 | 3497.1 | 0.28 | 172.5 W |
| NVIDIA A10G | 31.26 | 1185.2 | 0.28 | 110.2 W |
| NVIDIA L4 | 17.29 | 980.2 | 0.29 | 59.0 W |
A reasoning judge is the consensus endgame. Our consensus recipe needs two things at the referee seat: honesty (uncensored, so critiques aren't softened) and rigor (reasoning, so verdicts come with a visible chain of logic). This model is the only one we benchmark that has both. Put your Qwen and DeepSeek answers in front of it and it produces a reasoned, unhedged adjudication, the closest thing to a working code-review-bot personality in the open-source lineup. Vintage note: as a recent-generation Dolphin it inherits the line's gradual drift toward neutrality, the early Dolphins were blunter. For the judge role that's an acceptable trade; the reasoning trace matters more than maximum edge.
Fit and speed. The ~15GB floor means a 16GB card technically holds it, but tightly, long reasoning traces plus context will crowd the ceiling, so 24GB is the comfortable home. Speed lands mid-pack for its size: 121 tok/s peak, 109 on the H200 at 257W. Remember it pays the reasoning-token tax on top: on the A10G's 31 tok/s, a long adjudication takes a minute. If the judge is in your daily loop, feed it bandwidth.
Dolphin 3.0 R1 Mistral 24B: 121 tok/s peak, ~15GB floor: uncensored and reasoning-trained, which makes it our pick for the judge seat in a serious consensus setup. Give it 24GB for comfort and enough tok/s that its deliberation doesn't become your bottleneck.