NVIDIA B200, AI & Machine Learning Benchmarks & Specs

192GB · AI Score 78.0/100 · first-party measured on 12 AI workloads

78 AI Score ✓ Measured

Every number on this page is first-party: NVIDIA B200 was run on our pinned 12-workload AI suite on 2026-07-10, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) NVIDIA B200 delivers about 274.41 tokens/sec. Stepping up to Qwen3 32B it holds roughly 78.56 tok/s. The full Llama 3.3 70B still runs, at about 44.54 tok/s. For image generation, SDXL runs at 23.06 it/s, and FLUX.1-dev at 6.13 it/s. All 12 workloads fit in 192GB. There is no model in our suite this card has to turn down. NVIDIA B200 isn't a retail purchase for most people. It's rented by the hour. You can run this exact card on RunPod.

AI & Machine Learning benchmark results

Text Generation tok/s 123

LFM2.5-1.2B916.79
gemma-3-270m889.52
Llama 3.2 1B881.78
Qwen2.5-Coder-0.5B828.48
Qwen2.5-0.5B818.27
SmolLM2-135M815.3
Qwen1.5-0.5B699.73
Qwen3 0.6B599.63
Qwen3-0.6B596.73
LFM2.5-8B-A1B575.15
Qwen3-1.7B551.93
Qwen3 1.7B551.76
WorkloadResultTelemetryData
LFM2.5-1.2B916.79 tok/s
269 W33°CQ4_K_M
✓ Measured
gemma-3-270m889.52 tok/s
254 W28°CQ4_K_M
✓ Measured
Llama 3.2 1B881.78 tok/s
294 W34°CQ4_K_M
✓ Measured
Qwen2.5-Coder-0.5B828.48 tok/s
255 W29°CQ4_K_M
✓ Measured
Qwen2.5-0.5B818.27 tok/s
252 W29°CQ4_K_M
✓ Measured
SmolLM2-135M815.3 tok/s
252 W29°CQ4_K_M
✓ Measured
Qwen1.5-0.5B699.73 tok/s
267 W29°CQ4_K_M
✓ Measured
Qwen3 0.6B599.63 tok/s
238 W31°CQ4_K_M
✓ Measured
Qwen3-0.6B596.73 tok/s
266 W30°CQ4_K_M
✓ Measured
LFM2.5-8B-A1B575.15 tok/s
287 W31°CQ4_K_M
✓ Measured
Qwen3-1.7B551.93 tok/s
288 W32°CQ4_K_M
✓ Measured
Qwen3 1.7B551.76 tok/s
278 W33°CQ4_K_M
✓ Measured
gemma-3-1b519.19 tok/s
275 W30°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 1.5B456.7 tok/s
312 W33°CQ4_K_M
✓ Measured
Qwen2.5-Coder-1.5B455.35 tok/s
293 W32°CQ4_K_M
✓ Measured
Qwen2.5-1.5B454.69 tok/s
294 W30°CQ4_K_M
✓ Measured
Qwen2-1.5B454.55 tok/s
288 W31°CQ4_K_M
✓ Measured
Llama 3.2 3B422.09 tok/s
325 W35°CQ4_K_M
✓ Measured
Hermes-3-Llama-3.2-3B420.6 tok/s
331 W32°CQ4_K_M
✓ Measured
Llama-3.2-3B-Instruct-uncensored419.24 tok/s
320 W34°CQ4_K_M
✓ Measured
SmolLM3 3B400.73 tok/s
334 W35°CQ4_K_M
✓ Measured
SmolLM3-3B399.05 tok/s
322 W33°CQ4_K_M
✓ Measured
gemma-2-2b-it-abliterated392.32 tok/s
322 W32°CQ4_K_M
✓ Measured
gemma-2-2b391.94 tok/s
319 W32°CQ4_K_M
✓ Measured
Qwen2.5-Coder-3B388.17 tok/s
324 W33°CQ4_K_M
✓ Measured
Qwen2.5-3B388.14 tok/s
319 W32°CQ4_K_M
✓ Measured
Phi-4 Mini 3.8B375.74 tok/s
329 W36°CQ4_K_M
✓ Measured
Phi-4-mini375.06 tok/s
324 W34°CQ4_K_M
✓ Measured
AI21-Jamba-Reasoning-3B360.54 tok/s
332 W32°CQ4_K_M
✓ Measured
gpt-oss-20b351.42 tok/s
318 W32°CQ4_K_M
✓ Measured
phi-2348.72 tok/s
325 W32°CQ4_K_M
✓ Measured
Nemotron-3-Nano-30B-A3B345.93 tok/s
306 W32°CQ4_K_M
✓ Measured
DeepSeek-Coder-V2-Lite325.85 tok/s
295 W32°CQ4_K_M
✓ Measured
Phi-3.5-mini325 tok/s
328 W32°CQ4_K_M
✓ Measured
Qwen3 4B317.93 tok/s
3.1 GB peak288 W35°C1.11 tok/WQ4_K_M
✓ Measured
Qwen3-4B-Thinking-2507317.03 tok/s
341 W33°CQ4_K_M
✓ Measured
Qwen3-4B-Instruct-2507316.7 tok/s
349 W33°CQ4_K_M
✓ Measured
Qwen3-4B-Instruct-2507316.66 tok/s
348 W33°CQ4_K_M
✓ Measured
Llama-2-7B292 tok/s
391 W35°CQ4_K_M
✓ Measured
Gemma 3 4B291.82 tok/s
331 W35°CQ4_K_M
✓ Measured
Mistral 7B v0.3287.42 tok/s
324 W37°CQ4_K_M
✓ Measured
Mistral-7B-Instruct-v0.1286.89 tok/s
390 W36°CQ4_K_M
✓ Measured
Mistral-7B-Instruct-v0.3286.48 tok/s
412 W35°CQ4_K_M
✓ Measured
Mistral-7B-Instruct-v0.2286.37 tok/s
375 W35°CQ4_K_M
✓ Measured
Qwen2.5-Coder 7B277.98 tok/s
341 W38°CQ4_K_M
✓ Measured
Qwen3-Coder 30B A3B277.68 tok/s
282 W37°CQ4_K_M
✓ Measured
Qwen2.5-7B277.17 tok/s
387 W35°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 7B277.02 tok/s
299 W35°CQ4_K_M
✓ Measured
Qwen2.5-Coder-7B-Instruct-abliterated276.57 tok/s
391 W35°CQ4_K_M
✓ Measured
dolphin-2.9-llama3-8b274.48 tok/s
395 W35°CQ4_K_M
✓ Measured
L3-8B-Stheno-v3.2274.48 tok/s
382 W35°CQ4_K_M
✓ Measured
Llama 3.1 8B274.41 tok/s
5.1 GB peak373 W39°C0.74 tok/WQ4_K_M
✓ Measured
DarkIdol-Llama-3.1-8B-Instruct-1.2-Uncensored274.3 tok/s
398 W36°CQ4_K_M
✓ Measured
Meta-Llama-3.1-8B274.17 tok/s
397 W35°CQ4_K_M
✓ Measured
Dolphin 3.0 Llama 3.1 8B274.05 tok/s
363 W41°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill Llama 8B273.9 tok/s
314 W35°CQ4_K_M
✓ Measured
Dolphin X1 8B273.85 tok/s
327 W41°CQ4_K_M
✓ Measured
Qwen3-30B-A3B271.69 tok/s
300 W31°CQ4_K_M
✓ Measured
Qwen3 30B A3B270.84 tok/s
269 W32°CQ4_K_M
✓ Measured
Qwen3 8B253.45 tok/s
284 W35°CQ4_K_M
✓ Measured
Josiefied-Qwen3-8B-abliterated-v1253.28 tok/s
390 W35°CQ4_K_M
✓ Measured
Qwen3-8B253.14 tok/s
401 W36°CQ4_K_M
✓ Measured
DeepSeek-R1-0528-Qwen3-8B252.98 tok/s
382 W35°CQ4_K_M
✓ Measured
KAT-Coder-V2.5-Dev230.68 tok/s
296 W31°CQ4_K_M
✓ Measured
Qwen3 30B A3B (Q3_K_M)229.58 tok/s
313 W32°CQ3_K_M
✓ Measured
Ornith-1.0-9B226.2 tok/s
384 W35°CQ4_K_M
✓ Measured
Qwen-AgentWorld-35B-A3B221.78 tok/s
319 W32°CQ4_K_M
✓ Measured
Ornith-1.0-35B221.13 tok/s
305 W31°CQ4_K_M
✓ Measured
Dolphin X1 Trinity Nano 6B209.31 tok/s
272 W28°CQ4_K_M
✓ Measured
Qwen3-Coder-Next189.76 tok/s
291 W31°CQ4_K_M
✓ Measured
Qwen3-Next-80B-A3B-Thinking188.78 tok/s
288 W31°CQ4_K_M
✓ Measured
gemma-2-9b187.45 tok/s
414 W35°CQ4_K_M
✓ Measured
Qwen3-Coder-Next-abliterated186.57 tok/s
283 W31°CQ4_K_M
✓ Measured
Qwen3-Coder-Next186.39 tok/s
302 W31°CQ4_K_M
✓ Measured
Qwen3-Next-80B-A3B-Thinking186.03 tok/s
289 W31°CQ4_K_M
✓ Measured
NemoMix-Unleashed-12B183.93 tok/s
426 W36°CQ4_K_M
✓ Measured
Mistral-Nemo-Instruct-2407183.92 tok/s
427 W36°CQ4_K_M
✓ Measured
GLM-4.7-Flash183.82 tok/s
309 W31°CQ4_K_M
✓ Measured
Qwen3-Next-80B-A3B179.71 tok/s
296 W32°CQ4_K_M
✓ Measured
GLM-4.7-Flash-REAP-23B-A3B167.76 tok/s
307 W31°CQ4_K_M
✓ Measured
Phi-4 14B163.85 tok/s
365 W37°CQ4_K_M
✓ Measured
Qwen3 14B158.34 tok/s
316 W37°CQ4_K_M
✓ Measured
Hermes-4-14B158.22 tok/s
440 W36°CQ4_K_M
✓ Measured
Qwen3-14B158.19 tok/s
445 W36°CQ4_K_M
✓ Measured
Gemma 3 12B151.41 tok/s
291 W35°CQ4_K_M
✓ Measured
Qwen2.5-Coder 14B150.96 tok/s
9.1 GB peak399 W40°C0.38 tok/WQ4_K_M
✓ Measured
DeepSeek-R1 Distill 14B150.69 tok/s
331 W37°CQ4_K_M
✓ Measured
EVA-Qwen2.5-14B-v0.2150.67 tok/s
442 W36°CQ4_K_M
✓ Measured
Qwen2.5-Coder-14B-Instruct-abliterated150.67 tok/s
443 W36°CQ4_K_M
✓ Measured
Qwen2.5-14B150.57 tok/s
432 W36°CQ4_K_M
✓ Measured
Uncensored150.51 tok/s
433 W36°CQ4_K_M
✓ Measured
Gemma 4 12B146.36 tok/s
336 W35°CQ4_K_M
✓ Measured
StarCoder2 15B138.64 tok/s
345 W42°CQ4_K_M
✓ Measured
Phi-4 14B (Q3_K_M)135.21 tok/s
451 W36°CQ3_K_M
✓ Measured
Gemma 3 12B (Q3_K_M)125.31 tok/s
425 W35°CQ3_K_M
✓ Measured
DeepSeek-R1 Distill 14B (Q3_K_M)120.47 tok/s
433 W35°CQ3_K_M
✓ Measured
Mistral Small 24B114.04 tok/s
331 W38°CQ4_K_M
✓ Measured
Dolphin 3.0 R1 Mistral 24B113.95 tok/s
333 W44°CQ4_K_M
✓ Measured
Dolphin Mistral 24B Venice113.93 tok/s
365 W44°CQ4_K_M
✓ Measured
Devstral Small 24B113.91 tok/s
374 W44°CQ4_K_M
✓ Measured
Cydonia-24B-v4.3113.81 tok/s
476 W38°CQ4_K_M
✓ Measured
Dolphin-Mistral-24B-Venice-Edition113.8 tok/s
473 W38°CQ4_K_M
✓ Measured
Codestral 22B112.16 tok/s
366 W43°CQ4_K_M
✓ Measured
Codestral 22B (Q3_K_M)89.9 tok/s
475 W37°CQ3_K_M
✓ Measured
Mistral Small 24B (Q3_K_M)87.72 tok/s
483 W37°CQ3_K_M
✓ Measured
Gemma 3 27B86.08 tok/s
380 W43°CQ4_K_M
✓ Measured
Olmo-3.1-32B-Think79.75 tok/s
503 W38°CQ4_K_M
✓ Measured
Dolphin 2.9.1 Yi 1.5 34B79.35 tok/s
373 W44°CQ4_K_M
✓ Measured
Qwen3 32B78.56 tok/s
19 GB peak352 W42°C0.22 tok/WQ4_K_M
✓ Measured
DeepSeek-R1 Distill 32B78.55 tok/s
392 W44°CQ4_K_M
✓ Measured
Qwen2.5-32B78.55 tok/s
499 W38°CQ4_K_M
✓ Measured
Qwen2.5-Coder 32B78.53 tok/s
363 W44°CQ4_K_M
✓ Measured
DeepSeek-R1-Distill-Qwen-32B-abliterated78.48 tok/s
508 W38°CQ4_K_M
✓ Measured
QwQ 32B78.47 tok/s
349 W44°CQ4_K_M
✓ Measured
Qwen2.5-Coder 32B (Q3_K_M)60.24 tok/s
495 W37°CQ3_K_M
✓ Measured
Qwen2.5-72B45.34 tok/s
556 W42°CQ4_K_M
✓ Measured
Llama 3.3 70B44.54 tok/s
40.2 GB peak387 W47°C0.12 tok/WQ4_K_M
✓ Measured
DeepSeek-R1-Distill-Llama-70B44.49 tok/s
604 W41°CQ4_K_M
✓ Measured
Hermes-4-70B44.49 tok/s
583 W41°CQ4_K_M
✓ Measured
Llama-3.3-70B-Instruct-abliterated44.48 tok/s
575 W41°CQ4_K_M
✓ Measured
Meta-Llama-3.1-70B44.48 tok/s
594 W41°CQ4_K_M
✓ Measured
Laguna-XS-2.1✕ Won't fit needs ~26 GBVRAM-gated at this precision✓ Measured
Nanbeige4.2-3B✕ Won't fit needs ~4 GBVRAM-gated at this precision✓ Measured

Image Generation images/min 6

FLUX.1 Schnell81.41
Z-Image Turbo (1024px)62.47
Stable Diffusion XL46.12
Z-Image Turbo32.325
FLUX.1 dev13.136
Krea 2 Turbo9.5
WorkloadResultTelemetryData
FLUX.1 Schnell81.41 images/min
770 W43°C
✓ Measured
Z-Image Turbo (1024px)62.47 images/min
850 W45°C
✓ Measured
Stable Diffusion XL46.12 images/min
16.3 GB peak719 W44°C1.3 s/img
✓ Measured
Z-Image Turbo32.33 images/min
26.1 GB peak898 W50°C1.9 s/img
✓ Measured
FLUX.1 dev13.14 images/min
37 GB peak883 W51°C4.6 s/img
✓ Measured
Krea 2 Turbo9.5 images/min
801 W59°C
✓ Measured

Fine-Tuning train tok/s 4

SmolLM2 1.7B LoRA18064.7
TinyLlama 1.1B LoRA17993.5
Qwen2.5 1.5B LoRA15322
Qwen2.5 7B LoRA13722
WorkloadResultTelemetryData
SmolLM2 1.7B LoRA18064.7 train tok/s
435 W42°C
✓ Measured
TinyLlama 1.1B LoRA17993.5 train tok/s
363 W40°C
✓ Measured
Qwen2.5 1.5B LoRA15322 train tok/s
388 W40°C
✓ Measured
Qwen2.5 7B LoRA13722 train tok/s
692 W50°C
✓ Measured

LLM Serving serve tok/s 4

TinyLlama 1.1B served11562.2
SmolLM2 1.7B served9672.9
Qwen2.5 1.5B served9186.4
Qwen2.5 7B served5799.1
WorkloadResultTelemetryData
TinyLlama 1.1B served11562.2 serve tok/s
314 W33°C
✓ Measured
SmolLM2 1.7B served9672.9 serve tok/s
356 W34°C
✓ Measured
Qwen2.5 1.5B served9186.4 serve tok/s
332 W33°C
✓ Measured
Qwen2.5 7B served5799.1 serve tok/s
559 W37°C
✓ Measured

Image Editing images/min 2

WorkloadResultTelemetryData
FLUX.1 Kontext dev5.72 images/min
35.8 GB peak884 W53°C10.5 s/img
✓ Measured
Qwen-Image-Edit4.86 images/min
60.5 GB peak881 W52°C12.3 s/img
✓ Measured

Video Generation frames/s 2

WorkloadResultTelemetryData
LTX-Video (distilled)26.8 frames/s
60.5 GB peak730 W46°C3.6 s/clip
✓ Measured
Wan 2.2 5B (720p)1.88 frames/s
38 GB peak893 W51°C26 s/clip
✓ Measured

Depth Estimation images/min 2

WorkloadResultTelemetryData
Depth Anything V2 Small1137.49 images/min
190 W28°C
✓ Measured
Depth Anything V2 Large1039.47 images/min
192 W29°C
✓ Measured

Segmentation images/min 2

WorkloadResultTelemetryData
SAM ViT-Base1662.34 images/min
204 W29°C
✓ Measured
SAM ViT-Huge388.61 images/min
229 W32°C
✓ Measured

Background Removal images/min 1

WorkloadResultTelemetryData
BiRefNet1082.03 images/min
252 W31°C
✓ Measured

Upscaling images/min 1

WorkloadResultTelemetryData
Swin2SR 4x Upscaler11.97 images/min
282 W34°C
✓ Measured
How we measured this. Every result comes from our own pinned, reproducible AI suite, 12 workloads: the Qwen3-4B to Llama-70B LLM ladder (llama.cpp, Q4_K_M), SDXL / Z-Image / FLUX-dev generation, FLUX-Kontext / Qwen-Edit editing, and LTX / Wan video, run first-party on rented hardware with under 0.5% run-to-run variance. Peak VRAM, power draw, temperature and tokens-per-watt are captured per workload. “Won’t fit” rows are real data: where a model exceeds the card’s VRAM at the tested precision we record a hard gate rather than silently dropping to a smaller quant. Measured 2026-07-10 · harness 2.0.0.

NVIDIA B200 specifications

ArchitectureBlackwell
VRAM192GB HBM3e
Memory bus8192-bit
Memory bandwidth8 TB/s
TDP1000 W
ProcessTSMC 4NP
InterfaceSXM6
Release date2025-02-01
Launch MSRP$40,000

Verdict, nothing in our suite slows it down

NVIDIA B200 scores 78.0/100, #2 of 102. It ran all 12 workloads. Every figure here is our own measurement.

Relative performance: where the NVIDIA B200 lands

100% = this card, AI & Machine Learning headline metric (AI Score). #2 of 21 desktop cards in this vertical.

GPURelative%AI Score
NVIDIA B300
120%93.8
NVIDIA B200
100%78
NVIDIA B100
86%67
NVIDIA H100 NVL
86%67
NVIDIA GH200 Grace Hopper
84%65.6
NVIDIA H200
83%65

← All AI & Machine Learning GPU rankings

What this card can build

Whole-job timings, composed from our measured per-model results on this card.

WorkflowTimeEnergyBasis
50-image depth pass6 s0.15 Whmeasured
500-image masking run85 s4.9 Whmeasured
24-frame storyboard89 s12.85 Whall 2 stages measured
60-second AI short film1.9 min16.03 Whall 3 stages measured
6-panel comic page2.3 min23.05 Whall 3 stages measured
Character sheet, 12 poses2.9 min32.04 Whall 2 stages measured
10 short social clips5.6 min69.91 Whall 3 stages measured
Full codebase review6.6 min44.06 Whmeasured
40-product photo shoot8.4 min113.44 Whall 2 stages measured
20 long-form articles10.5 min67.6 Whmeasured
40-product shoot, start to finish12.1 min129.28 Whall 4 stages measured
200-product catalogue cutout17.2 min79.19 Whall 2 stages measured
100-photo restoration batch17.8 min257.65 Whmeasured
100-photo restore and enlarge26.4 min296.86 Whall 2 stages measured

Rent or buy?

This card is $40,000 to buy. The cheapest listed rate on RunPod is $5.980/hour, but that is the floor: we budget $7.176/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 5,574 GPU-hours. Below it you are paying for idle silicon.

How you would use itGPU-hours a yearRental cost a yearTime to break even
2 hours a day, hobby730$5,2387.6 years
8 hours a day, working on it2,920$20,9541.9 years
24/7, always-on agent8,760$62,8627.6 months

At hobby usage this card is very unlikely to pay for itself before it is superseded. Rent it. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.

Rental price

$5.980/hr+1.5% since 2026-08-14low $5.890 · high $5.980

Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.