Mistral 7B v0.3 · 11 GPUs measured first-party · llama.cpp Q4_K_M · Updated October 2026

What GPU Do You Need for Mistral 7B v0.3?

Mistral 7B is the model that made local LLMs real, the 2023 release that proved small open models could punch far above their weight. We measured it on 11 GPUs (llama.cpp, Q4_K_M): 299 tok/s on the B300, ~5GB peak VRAM. It remains one of the fastest 7Bs on our bench, and, honestly, a piece of history more than a current recommendation.

Benchmarked weights: bartowski/Mistral-7B-Instruct-v0.3-GGUF

Fastest we measured
NVIDIA B300

NVIDIA B300

298.9 tok/s on Mistral 7B v0.3, the ceiling. Measured on our bench. 288GB of VRAM, $40,000 at launch.

Pros
  • 298.9 tok/s on Mistral 7B v0.3
  • 288GB, clears the Mistral 7B v0.3 floor
  • Rentable by the hour rather than bought
Cons
  • 1400W board rating
  • Datacenter or workstation hardware, not a retail purchase
Cheapest card that runs it
NVIDIA RTX PRO 6000 Blackwell Workstation Edition

NVIDIA RTX PRO 6000 Blackwell Workstation Edition

247.3 tok/s on Mistral 7B v0.3, lowest launch price that still fits. Measured on our bench. 96GB of VRAM, $8,565 at launch.

Pros
  • 247.3 tok/s on Mistral 7B v0.3
  • 96GB, clears the Mistral 7B v0.3 floor
  • Rentable by the hour rather than bought
Cons
  • 600W board rating
  • Datacenter or workstation hardware, not a retail purchase
298.9tok/s
Fastest: NVIDIA B300
measured
11
Cards that run Mistral 7B v0.3
of 11 we have data for
0
Cards that can't run it at all
published as hard gates, not omissions
647%
Fastest vs slowest that fits
298.9 vs 40.01 tok/s

What GPU Do You Need for Mistral 7B v0.3?, tok/s by GPU

NVIDIA B300
298.9 tok/s
NVIDIA B200
287.4 tok/s
NVIDIA H200
278.5 tok/s
NVIDIA H100 80GB HBM3
275.9 tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
247.3 tok/s
NVIDIA A100 80GB SXM4
171.9 tok/s
NVIDIA A100 40GB SXM4
165.9 tok/s
NVIDIA L40S
144.7 tok/s
NVIDIA A10G
92.82 tok/s
NVIDIA L4
54.02 tok/s
NVIDIA T4
40.01 tok/s

Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.

Efficiency: tok/s per 100W drawn

NVIDIA H200
264.23 tok/s / 100W
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
141.09 tok/s / 100W
NVIDIA A100 80GB SXM4
133.89 tok/s / 100W
NVIDIA H100 80GB HBM3
117.63 tok/s / 100W
NVIDIA B300
95.42 tok/s / 100W
NVIDIA A100 40GB SXM4
91 tok/s / 100W
NVIDIA B200
88.63 tok/s / 100W
NVIDIA L4
84.94 tok/s / 100W
NVIDIA A10G
74.43 tok/s / 100W
NVIDIA L40S
70.67 tok/s / 100W
NVIDIA T4
63.01 tok/s / 100W

Power is the average pulled during the run, sampled at 1Hz. The fastest card is often not the one here, and for anything left running this is the number that shows up on the bill.

Value: tok/s per $1,000 of MSRP

NVIDIA A10G
33.15 tok/s / $1k
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
28.88 tok/s / $1k
NVIDIA L4
21.61 tok/s / $1k
NVIDIA L40S
19.3 tok/s / $1k
NVIDIA T4
17.4 tok/s / $1k
NVIDIA A100 40GB SXM4
13.83 tok/s / $1k
NVIDIA A100 80GB SXM4
10.11 tok/s / $1k
NVIDIA H100 80GB HBM3
9.2 tok/s / $1k
NVIDIA H200
8.98 tok/s / $1k
NVIDIA B300
7.47 tok/s / $1k
NVIDIA B200
7.19 tok/s / $1k

Launch price, not street price, so it ages. A speed leaderboard always crowns the most expensive card; this is the counterweight.

Mistral 7B v0.3. Measured generation speed by GPU

NVIDIA B300298.9
NVIDIA B200287.4
NVIDIA H200278.5
NVIDIA H100 80GB HBM3275.9
NVIDIA RTX PRO 6000 Blackwell Workstation Edition247.3
NVIDIA A100 80GB SXM4171.9
NVIDIA A100 40GB SXM4165.9
NVIDIA L40S144.7
NVIDIA A10G92.82
NVIDIA L454.02
NVIDIA T440.01
GPUtok/sPrompt t/stok/WAvg power
NVIDIA B300298.95478.50.95313.3 W
NVIDIA B200287.410045.10.89324.3 W
NVIDIA H200278.590312.64105.4 W
NVIDIA H100 80GB HBM3275.99353.51.18234.5 W
NVIDIA RTX PRO 6000 Blackwell Workstation Edition247.311995.31.41175.3 W
NVIDIA A100 80GB SXM4171.94516.51.34128.4 W
NVIDIA A100 40GB SXM4165.94365.50.91182.3 W
NVIDIA L40S144.710048.80.71204.8 W
NVIDIA A10G92.823201.40.74124.7 W
NVIDIA L454.023032.80.8563.6 W
NVIDIA T440.011178.90.6363.5 W

Our take: respect it, don't run it. Two years is a geological age in open models, and it shows. Today our floor for serious local work is 12-14B, with the 27-32B range as the sweet spot: and even inside the small tier, Qwen3 8B, the DeepSeek 7B distill and Phi-4 Mini all outclass Mistral 7B's output while costing the same VRAM. What survives is everything around the model: the enormous fine-tune ecosystem built on this base (the 8B Dolphins in our own database included), its permissive license, and its status as the default 'known quantity' base for custom tunes.

Still quick, for what it's worth. 299 tok/s peak, 278 on the H200 at a superb 2.64 tok/W, 40 tok/s even on a T4: the architecture's efficiency was always its magic, and the ~5GB floor undercuts most of its successors. If you do have a reason to run it (a fine-tune you love, a legacy pipeline, minimal hardware), it costs almost nothing to host. Just don't mistake fast for good in 2026.

How it compares. H100 80GB HBM3: Mistral 7B v0.3 275.9 tok/s, Mistral-7B-Instruct-v0.1 276.3, Mistral-7B-Instruct-v0.3 276.6, Mistral-7B-Instruct-v0.2 276.7 (7B), Qwen3 30B A3B 283.8 (31B). All 4 beat Mistral 7B v0.3 here. Qwen3 30B A3B is a mixture-of-experts, so per token it computes only a slice of its size.

Cost on a rented GPU. 1M generated tokens of Mistral 7B v0.3: $0.79 on a A100 40GB SXM4 ($0.47/hr, 100 min), $6.45 on a B300 ($6.94/hr, 56 min, 8.2x the cost).

Mistral 7B v0.3: cost per 1M generated tokens on rented GPUs

NVIDIA A100 40GB SXM4$0.47/hr
NVIDIA T4$0.14/hr
NVIDIA RTX PRO 6000 Blackwell Workstation Edition$1.08/hr
NVIDIA L40S$0.79/hr
NVIDIA A100 80GB SXM4$0.95/hr
NVIDIA H100 80GB HBM3$2.14/hr
NVIDIA L4$0.44/hr
NVIDIA H200$3.59/hr
NVIDIA B200$5.98/hr
NVIDIA B300$6.94/hr
GPUCheapest rateSpeed (tok/s)Cost per 1M generated tokens
NVIDIA A100 40GB SXM4$0.47/hr165.9$0.79
NVIDIA T4$0.14/hr40.01$0.94
NVIDIA RTX PRO 6000 Blackwell Workstation Edition$1.08/hr247.3$1.21
NVIDIA L40S$0.79/hr144.7$1.52
NVIDIA A100 80GB SXM4$0.95/hr171.9$1.53
NVIDIA H100 80GB HBM3$2.14/hr275.9$2.15
NVIDIA L4$0.44/hr54.02$2.26
NVIDIA H200$3.59/hr278.5$3.58
NVIDIA B200$5.98/hr287.4$5.78
NVIDIA B300$6.94/hr298.9$6.45

Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.

Speed tiers for Mistral 7B v0.3. 30+ tok/s: 11 (B300, B200, H200). 30 tok/s is roughly where replies outpace reading.

Reading your prompt. Before Mistral 7B v0.3 writes anything it reads the input: 11995.3 tok/s on the RTX PRO 6000 Blackwell Workstation Edition (0.3s for a 4,000-token prompt), 1178.9 on the T4 (3.4s). Long documents and big code files feel this number more than the generation speed.

VRAM for Mistral 7B v0.3. Measured peak 4.7GB, so 8GB is the smallest common card size; smallest card it ran on: T4 (16GB). With long context: Q4_K_M 5GB (tested), Q2_K 4GB, Q3_K_M 5GB, Q5_K_M 6GB, Q6_K 7GB.

Power on Mistral 7B v0.3. Most efficient: RTX PRO 6000 Blackwell Workstation Edition, 175W, 0.20 kWh per 1M generated tokens. Hungriest: B200, 324W, 0.31 kWh. At $0.15/kWh: $0.030 per 1M generated tokens.

Our verdict

Mistral 7B: 299 tok/s peak, ~5GB floor, historically important and efficient to this day, but outclassed at its own size by 2025-era models and far below our 12-14B minimum for serious work. Run its descendants, or run it as a base to tune; as a daily model its moment has passed.

FAQ

Is Mistral 7B still worth using?
As a daily assistant, no, same-size successors (Qwen3 8B, Phi-4 Mini, the DeepSeek 7B distill) answer clearly better on identical hardware. As a fine-tuning base with a massive ecosystem, it's still a legitimate pick.
What hardware does it need?
Almost any GPU: ~5GB measured peak at Q4_K_M, 299 tok/s at the top, 40 tok/s on a 2018 T4. Its efficiency (2.64 tok/W on the H200) is still excellent.
What's the minimum model size you'd recommend today?
Our working rule: 12-14B minimum for dependable output, 27-32B as the sweet spot. The 7B tier is for pipelines, completions and constrained hardware, not primary assistants.
Why does this model matter historically?
Its 2023 release proved a well-trained 7B could rival much larger models, effectively launching the local-LLM era. Half the small-model ecosystem, including several Dolphins we benchmark, stands on this base.
What should I run instead on the same hardware?
Qwen3 8B (~6GB, 261 tok/s peak in our data) for general chat, the DeepSeek 7B distill for reasoning, or Qwen2.5-Coder 7B for completions, each a straight upgrade at the same VRAM class.