Mistral 7B v0.3 · 11 GPUs measured first-party · llama.cpp Q4_K_M · Updated July 2026
Mistral 7B is the model that made local LLMs real, the 2023 release that proved small open models could punch far above their weight. We measured it on 11 GPUs (llama.cpp, Q4_K_M): 299 tok/s on the B300, ~5GB peak VRAM. It remains one of the fastest 7Bs on our bench, and, honestly, a piece of history more than a current recommendation.
Benchmarked weights: bartowski/Mistral-7B-Instruct-v0.3-GGUF
What GPU Do You Need for Mistral 7B v0.3?, tok/s, fastest 11
Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.
Mistral 7B v0.3. Measured generation speed by GPU
| GPU | tok/s | Prompt t/s | tok/W | Avg power |
|---|---|---|---|---|
| NVIDIA B300 | 298.95 | 5478.5 | 0.95 | 313.3 W |
| NVIDIA B200 | 287.42 | 10045.1 | 0.89 | 324.3 W |
| NVIDIA H200 | 278.5 | 9031 | 2.64 | 105.4 W |
| NVIDIA H100 80GB HBM3 | 275.85 | 9353.5 | 1.18 | 234.5 W |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 247.33 | 11995.3 | 1.41 | 175.3 W |
| NVIDIA A100 80GB SXM4 | 171.92 | 4516.5 | 1.34 | 128.4 W |
| NVIDIA A100 40GB SXM4 | 167.21 | 4334.4 | 1.27 | 131.2 W |
| NVIDIA L40S | 144.97 | 10241.1 | 0.86 | 169.5 W |
| NVIDIA A10G | 99.35 | 3897.7 | 0.8 | 123.9 W |
| NVIDIA L4 | 53.28 | 2870.6 | 0.95 | 56.2 W |
| NVIDIA T4 | 40.2 | 1226.4 | 0.72 | 56.1 W |
Our take: respect it, don't run it. Two years is a geological age in open models, and it shows. Today our floor for serious local work is 12-14B, with the 27-32B range as the sweet spot: and even inside the small tier, Qwen3 8B, the DeepSeek 7B distill and Phi-4 Mini all outclass Mistral 7B's output while costing the same VRAM. What survives is everything around the model: the enormous fine-tune ecosystem built on this base (the 8B Dolphins in our own database included), its permissive license, and its status as the default 'known quantity' base for custom tunes.
Still quick, for what it's worth. 299 tok/s peak, 278 on the H200 at a superb 2.64 tok/W, 40 tok/s even on a T4: the architecture's efficiency was always its magic, and the ~5GB floor undercuts most of its successors. If you do have a reason to run it (a fine-tune you love, a legacy pipeline, minimal hardware), it costs almost nothing to host. Just don't mistake fast for good in 2026.
Mistral 7B: 299 tok/s peak, ~5GB floor, historically important and efficient to this day, but outclassed at its own size by 2025-era models and far below our 12-14B minimum for serious work. Run its descendants, or run it as a base to tune; as a daily model its moment has passed.