Dolphin Mistral 24B Venice · 10 GPUs measured first-party · llama.cpp Q4_K_M · Updated July 2026
Dolphin Mistral 24B Venice Edition is the direct-answer counterpart to the R1 Dolphin: the same uncensored Mistral 24B foundation, tuned for straight responses instead of reasoning traces. Measured on 10 GPUs (llama.cpp, Q4_K_M): 121 tok/s on the B300, ~15GB peak, and a standout efficiency run on the H200: 109 tok/s at just 110W.
Benchmarked weights: dphn/Dolphin-Mistral-24B-Venice-Edition-GGUF
What GPU Do You Need for Dolphin Mistral 24B Venice?, tok/s, fastest 10
Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.
Dolphin Mistral 24B Venice. Measured generation speed by GPU
| GPU | tok/s | Prompt t/s | tok/W | Avg power |
|---|---|---|---|---|
| NVIDIA B300 | 120.93 | 1951.9 | 0.37 | 330.8 W |
| NVIDIA B200 | 113.93 | 3816.4 | 0.31 | 364.9 W |
| NVIDIA H200 | 109.14 | 3417.3 | 0.99 | 110.0 W |
| NVIDIA H100 80GB HBM3 | 109.13 | 3425.3 | 0.57 | 191.3 W |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 93.67 | 4927.8 | 0.6 | 157.4 W |
| NVIDIA A100 40GB SXM4 | 62.98 | 1638.5 | 0.46 | 137.7 W |
| NVIDIA A100 80GB SXM4 | 62.72 | 1656.3 | 0.39 | 161.4 W |
| NVIDIA L40S | 48.31 | 3386.6 | 0.29 | 164.7 W |
| NVIDIA A10G | 33.01 | 1492.9 | 0.23 | 141.9 W |
| NVIDIA L4 | 16.89 | 963.2 | 0.3 | 57.2 W |
The fast half of the uncensored 24B pair. Venice skips the thinking-token tax: every generated token is answer, which makes it the better daily chat model of the two Mistral Dolphins, same brain size, no deliberation overhead. It's also become a privacy-community favorite (the Venice name comes from the private-AI platform that commissioned it), and running it locally is the logical conclusion of that: an unaligned, capable 24B that never leaves your machine. Set expectations accordingly: today's Dolphins, Venice included, are more neutral than the early generation was. The line has softened over the years even while staying well clear of aligned-model filtering. For most private local use that middle ground is exactly right.
The efficiency headline. The H200 row deserves attention: 109 tok/s at 110.0W measured, 0.99 tok/W, nearly triple the efficiency of the B300 that beats it by 11%. That's the widest Hopper-vs-Blackwell efficiency gap in our 24B data. On owned hardware, the ~15GB floor makes 16GB cards a snug fit and 24GB cards comfortable; at 33 tok/s even an A10G delivers usable chat.
Dolphin Mistral 24B Venice: 121 tok/s peak, ~15GB floor, no reasoning overhead, the fast, private, uncensored daily driver of the Dolphin line. If the R1 Dolphin is your judge, this is your workhorse; together they're the whole uncensored mid-size stack.