Mistral 7B v0.3 · 11 GPUs measured first-party · llama.cpp Q4_K_M · Updated October 2026
Mistral 7B is the model that made local LLMs real, the 2023 release that proved small open models could punch far above their weight. We measured it on 11 GPUs (llama.cpp, Q4_K_M): 299 tok/s on the B300, ~5GB peak VRAM. It remains one of the fastest 7Bs on our bench, and, honestly, a piece of history more than a current recommendation.
Benchmarked weights: bartowski/Mistral-7B-Instruct-v0.3-GGUF

298.9 tok/s on Mistral 7B v0.3, the ceiling. Measured on our bench. 288GB of VRAM, $40,000 at launch.

247.3 tok/s on Mistral 7B v0.3, lowest launch price that still fits. Measured on our bench. 96GB of VRAM, $8,565 at launch.
What GPU Do You Need for Mistral 7B v0.3?, tok/s by GPU
Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.
Efficiency: tok/s per 100W drawn
Power is the average pulled during the run, sampled at 1Hz. The fastest card is often not the one here, and for anything left running this is the number that shows up on the bill.
Value: tok/s per $1,000 of MSRP
Launch price, not street price, so it ages. A speed leaderboard always crowns the most expensive card; this is the counterweight.
Mistral 7B v0.3. Measured generation speed by GPU
| GPU | tok/s | Prompt t/s | tok/W | Avg power |
|---|---|---|---|---|
| NVIDIA B300 | 298.9 | 5478.5 | 0.95 | 313.3 W |
| NVIDIA B200 | 287.4 | 10045.1 | 0.89 | 324.3 W |
| NVIDIA H200 | 278.5 | 9031 | 2.64 | 105.4 W |
| NVIDIA H100 80GB HBM3 | 275.9 | 9353.5 | 1.18 | 234.5 W |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 247.3 | 11995.3 | 1.41 | 175.3 W |
| NVIDIA A100 80GB SXM4 | 171.9 | 4516.5 | 1.34 | 128.4 W |
| NVIDIA A100 40GB SXM4 | 165.9 | 4365.5 | 0.91 | 182.3 W |
| NVIDIA L40S | 144.7 | 10048.8 | 0.71 | 204.8 W |
| NVIDIA A10G | 92.82 | 3201.4 | 0.74 | 124.7 W |
| NVIDIA L4 | 54.02 | 3032.8 | 0.85 | 63.6 W |
| NVIDIA T4 | 40.01 | 1178.9 | 0.63 | 63.5 W |
Our take: respect it, don't run it. Two years is a geological age in open models, and it shows. Today our floor for serious local work is 12-14B, with the 27-32B range as the sweet spot: and even inside the small tier, Qwen3 8B, the DeepSeek 7B distill and Phi-4 Mini all outclass Mistral 7B's output while costing the same VRAM. What survives is everything around the model: the enormous fine-tune ecosystem built on this base (the 8B Dolphins in our own database included), its permissive license, and its status as the default 'known quantity' base for custom tunes.
Still quick, for what it's worth. 299 tok/s peak, 278 on the H200 at a superb 2.64 tok/W, 40 tok/s even on a T4: the architecture's efficiency was always its magic, and the ~5GB floor undercuts most of its successors. If you do have a reason to run it (a fine-tune you love, a legacy pipeline, minimal hardware), it costs almost nothing to host. Just don't mistake fast for good in 2026.
How it compares. H100 80GB HBM3: Mistral 7B v0.3 275.9 tok/s, Mistral-7B-Instruct-v0.1 276.3, Mistral-7B-Instruct-v0.3 276.6, Mistral-7B-Instruct-v0.2 276.7 (7B), Qwen3 30B A3B 283.8 (31B). All 4 beat Mistral 7B v0.3 here. Qwen3 30B A3B is a mixture-of-experts, so per token it computes only a slice of its size.
Cost on a rented GPU. 1M generated tokens of Mistral 7B v0.3: $0.79 on a A100 40GB SXM4 ($0.47/hr, 100 min), $6.45 on a B300 ($6.94/hr, 56 min, 8.2x the cost).
Mistral 7B v0.3: cost per 1M generated tokens on rented GPUs
| GPU | Cheapest rate | Speed (tok/s) | Cost per 1M generated tokens |
|---|---|---|---|
| NVIDIA A100 40GB SXM4 | $0.47/hr | 165.9 | $0.79 |
| NVIDIA T4 | $0.14/hr | 40.01 | $0.94 |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | $1.08/hr | 247.3 | $1.21 |
| NVIDIA L40S | $0.79/hr | 144.7 | $1.52 |
| NVIDIA A100 80GB SXM4 | $0.95/hr | 171.9 | $1.53 |
| NVIDIA H100 80GB HBM3 | $2.14/hr | 275.9 | $2.15 |
| NVIDIA L4 | $0.44/hr | 54.02 | $2.26 |
| NVIDIA H200 | $3.59/hr | 278.5 | $3.58 |
| NVIDIA B200 | $5.98/hr | 287.4 | $5.78 |
| NVIDIA B300 | $6.94/hr | 298.9 | $6.45 |
Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.
Speed tiers for Mistral 7B v0.3. 30+ tok/s: 11 (B300, B200, H200). 30 tok/s is roughly where replies outpace reading.
Reading your prompt. Before Mistral 7B v0.3 writes anything it reads the input: 11995.3 tok/s on the RTX PRO 6000 Blackwell Workstation Edition (0.3s for a 4,000-token prompt), 1178.9 on the T4 (3.4s). Long documents and big code files feel this number more than the generation speed.
VRAM for Mistral 7B v0.3. Measured peak 4.7GB, so 8GB is the smallest common card size; smallest card it ran on: T4 (16GB). With long context: Q4_K_M 5GB (tested), Q2_K 4GB, Q3_K_M 5GB, Q5_K_M 6GB, Q6_K 7GB.
Power on Mistral 7B v0.3. Most efficient: RTX PRO 6000 Blackwell Workstation Edition, 175W, 0.20 kWh per 1M generated tokens. Hungriest: B200, 324W, 0.31 kWh. At $0.15/kWh: $0.030 per 1M generated tokens.
Mistral 7B: 299 tok/s peak, ~5GB floor, historically important and efficient to this day, but outclassed at its own size by 2025-era models and far below our 12-14B minimum for serious work. Run its descendants, or run it as a base to tune; as a daily model its moment has passed.