NVIDIA B200 vs NVIDIA B300, AI & Machine Learning Comparison

NVIDIA B200
NVIDIA B200
vs
NVIDIA B300
NVIDIA B300

NVIDIA B300 wins 135 of 142 benchmarks, averaging 12.4% faster.

Both cards were measured first-party on our bench, same suite, same test rig.

What the numbers say

The gap is widest in Swin2SR 4x Upscaler, where NVIDIA B300 leads by 387% (11.97 vs 58.29 images/min); the closest fight is Nemotron-3-Nano-30B-A3B (0% apart).

Benchmark results head-to-head

BenchmarkNVIDIA B200NVIDIA B300Difference
Qwen3 4B tok/s317.93333.34-5%
Llama 3.1 8B tok/s274.41287.23-4%
Qwen2.5-Coder 14B tok/s150.96158.49-5%
Qwen3 32B tok/s78.5683.68-6%
Llama 3.3 70B tok/s44.5447.97-7%
Stable Diffusion XL images/min46.1229.2+58%
Z-Image Turbo images/min32.32539.675-19%
FLUX.1 dev images/min13.13620.807-37%
FLUX.1 Kontext dev images/min5.72110.329-45%
Qwen-Image-Edit images/min4.868.14-40%
LTX-Video (distilled) frames/s26.831.87-16%
Wan 2.2 5B (720p) frames/s1.882.94-36%
DeepSeek-R1 Distill Llama 8B tok/s273.9283.77-3%
DeepSeek-R1 Distill 1.5B tok/s456.7562.14-19%
DeepSeek-R1 Distill 14B tok/s150.69156.33-4%
DeepSeek-R1 Distill 7B tok/s277.02284.77-3%
Gemma 3 12B tok/s151.41161.27-6%
Gemma 3 4B tok/s291.82301.07-3%
Gemma 4 12B tok/s146.36155.96-6%
Llama 3.2 1B tok/s881.78889.19-1%
Llama 3.2 3B tok/s422.09432.24-2%
Mistral 7B v0.3 tok/s287.42298.95-4%
Mistral Small 24B tok/s114.04119.3-4%
Phi-4 14B tok/s163.85177.41-8%
Phi-4 Mini 3.8B tok/s375.74398.16-6%
Qwen2.5-Coder 7B tok/s277.98286.5-3%
Qwen3 0.6B tok/s599.63718.58-17%
Qwen3 1.7B tok/s551.76555.1-1%
Qwen3 14B tok/s158.34165.44-4%
Qwen3 30B A3B tok/s270.84284.66-5%
Qwen3 8B tok/s253.45261.4-3%
SmolLM3 3B tok/s400.73411.25-3%
Codestral 22B tok/s112.16119.81-6%
DeepSeek-R1 Distill 32B tok/s78.5582.9-5%
Devstral Small 24B tok/s113.91121.18-6%
Dolphin 2.9.1 Yi 1.5 34B tok/s79.3585.15-7%
Dolphin Mistral 24B Venice tok/s113.93120.93-6%
Dolphin X1 8B tok/s273.85287.02-5%
Dolphin 3.0 Llama 3.1 8B tok/s274.05285.88-4%
Dolphin 3.0 R1 Mistral 24B tok/s113.95120.95-6%
Gemma 3 27B tok/s86.0893.44-8%
Qwen2.5-Coder 32B tok/s78.5382.9-5%
Qwen3-Coder 30B A3B tok/s277.68284.88-3%
QwQ 32B tok/s78.4782.91-5%
StarCoder2 15B tok/s138.64144.05-4%
FLUX.1 Schnell images/min81.4159.54+37%
Z-Image Turbo (1024px) images/min62.4748.63+28%
BiRefNet images/min1082.031064.73+2%
Depth Anything V2 Large images/min1039.471069.95-3%
Depth Anything V2 Small images/min1137.491491.6-24%
SAM ViT-Base images/min1662.341719.59-3%
SAM ViT-Huge images/min388.61399.37-3%
Swin2SR 4x Upscaler images/min11.9758.29-79%
Qwen2.5 1.5B LoRA train tok/s1532222351.1-31%
Qwen2.5 7B LoRA train tok/s1372214323.1-4%
SmolLM2 1.7B LoRA train tok/s18064.726114.3-31%
TinyLlama 1.1B LoRA train tok/s17993.526938.2-33%
AI21-Jamba-Reasoning-3B tok/s360.54380.88-5%
Olmo-3.1-32B-Think tok/s79.7587.85-9%
Codestral 22B (Q3_K_M) tok/s89.9107.98-17%
Dolphin-Mistral-24B-Venice-Edition tok/s113.8121.87-7%
DarkIdol-Llama-3.1-8B-Instruct-1.2-Uncensored tok/s274.3292.55-6%
DeepSeek-Coder-V2-Lite tok/s325.85340.53-4%
DeepSeek-R1-0528-Qwen3-8B tok/s252.98270.35-6%
DeepSeek-R1-Distill-Llama-70B tok/s44.4948.11-8%
DeepSeek-R1 Distill 14B (Q3_K_M) tok/s120.47144.74-17%
DeepSeek-R1-Distill-Qwen-32B-abliterated tok/s78.4883.19-6%
dolphin-2.9-llama3-8b tok/s274.48290.47-6%
Dolphin X1 Trinity Nano 6B tok/s209.31257.93-19%
EVA-Qwen2.5-14B-v0.2 tok/s150.67159.46-6%
gemma-2-2b-it-abliterated tok/s392.32418.18-6%
gemma-2-2b tok/s391.94417.84-6%
gemma-2-9b tok/s187.45202.22-7%
Gemma 3 12B (Q3_K_M) tok/s125.31149.61-16%
gemma-3-1b tok/s519.19533.32-3%
gemma-3-270m tok/s889.52926.67-4%
GLM-4.7-Flash-REAP-23B-A3B tok/s167.76178.49-6%
GLM-4.7-Flash tok/s183.82196.52-6%
Josiefied-Qwen3-8B-abliterated-v1 tok/s253.28270.09-6%
gpt-oss-20b tok/s351.42355.56-1%
Hermes-3-Llama-3.2-3B tok/s420.6446.78-6%
Hermes-4-70B tok/s44.4948.1-8%
SmolLM3-3B tok/s399.05420.3-5%
Qwen3-Coder-Next-abliterated tok/s186.57199.12-6%
KAT-Coder-V2.5-Dev tok/s230.68239.66-4%
L3-8B-Stheno-v3.2 tok/s274.48293-6%
Laguna-XS-2.1 tok/s00n/a
LFM2.5-1.2B tok/s916.79955.13-4%
LFM2.5-8B-A1B tok/s575.15596.16-4%
Llama-2-7B tok/s292315.9-8%
Llama-3.2-3B-Instruct-uncensored tok/s419.24448.94-7%
Llama-3.3-70B-Instruct-abliterated tok/s44.4848.1-8%
Meta-Llama-3.1-70B tok/s44.4848.12-8%
Meta-Llama-3.1-8B tok/s274.17293-6%
Phi-4-mini tok/s375.06409.58-8%
Mistral-7B-Instruct-v0.1 tok/s286.89306.44-6%
Mistral-7B-Instruct-v0.2 tok/s286.37306.79-7%
Mistral-7B-Instruct-v0.3 tok/s286.48306.46-7%
Mistral-Nemo-Instruct-2407 tok/s183.92196.92-7%
Mistral Small 24B (Q3_K_M) tok/s87.72108.37-19%
Nanbeige4.2-3B tok/s00n/a
NemoMix-Unleashed-12B tok/s183.93197.06-7%
Nemotron-3-Nano-30B-A3B tok/s345.93345.310%
Hermes-4-14B tok/s158.22169.51-7%
Ornith-1.0-35B tok/s221.13232.36-5%
Ornith-1.0-9B tok/s226.2240.67-6%
phi-2 tok/s348.72356.65-2%
Phi-3.5-mini tok/s325352.25-8%
Phi-4 14B (Q3_K_M) tok/s135.21165.06-18%
Qwen-AgentWorld-35B-A3B tok/s221.78233.04-5%
Qwen3-0.6B tok/s596.73742.64-20%
Qwen3-1.7B tok/s551.93575.97-4%
Qwen3-14B tok/s158.19169.52-7%
Qwen3-30B-A3B tok/s271.69291.82-7%
Qwen3-4B-Instruct-2507 tok/s316.66339.78-7%
Qwen3-8B tok/s253.14270.35-6%
Qwen3-Coder-Next tok/s186.39196.34-5%
Qwen3-Next-80B-A3B-Thinking tok/s186.03194.98-5%
Qwen1.5-0.5B tok/s699.73894.84-22%
Qwen2-1.5B tok/s454.55583.79-22%
Qwen2.5-0.5B tok/s818.27843.4-3%
Qwen2.5-1.5B tok/s454.69584.96-22%
Uncensored tok/s150.51159.48-6%
Qwen2.5-14B tok/s150.57159.33-5%
Qwen2.5-32B tok/s78.5583.41-6%
Qwen2.5-3B tok/s388.14409.56-5%
Qwen2.5-72B tok/s45.3448.28-6%
Qwen2.5-7B tok/s277.17293.63-6%
Qwen2.5-Coder-0.5B tok/s828.48843.44-2%
Qwen2.5-Coder-1.5B tok/s455.35584.5-22%
Qwen2.5-Coder-14B-Instruct-abliterated tok/s150.67159.5-6%
Qwen2.5-Coder 32B (Q3_K_M) tok/s60.2474.28-19%
Qwen2.5-Coder-3B tok/s388.17410.06-5%
Qwen2.5-Coder-7B-Instruct-abliterated tok/s276.57292.21-5%
Qwen3 30B A3B (Q3_K_M) tok/s229.58268.71-15%
Qwen3-4B-Instruct-2507 tok/s316.7339.72-7%
Qwen3-4B-Thinking-2507 tok/s317.03339.74-7%
Qwen3-Coder-Next tok/s189.76199.39-5%
Qwen3-Next-80B-A3B-Thinking tok/s188.78200.55-6%
Qwen3-Next-80B-A3B tok/s179.71191.32-6%
SmolLM2-135M tok/s815.3855.87-5%
Cydonia-24B-v4.3 tok/s113.81121.9-7%

Whole-job comparison

How long each card takes to finish a complete pipeline, not just one model. NVIDIA B200 is faster on 2 of 14; NVIDIA B300 on 12.

WorkflowNVIDIA B200NVIDIA B300DifferenceCost per run
50-image depth pass6 s5 sNVIDIA B300 1.22x faster$0.009 vs $0.009
500-image masking run85 s78 sNVIDIA B300 1.10x faster$0.141 vs $0.150
24-frame storyboard89 s1.5 minNVIDIA B200 1.01x faster$0.149 vs $0.174
60-second AI short film1.9 min2.2 minNVIDIA B200 1.17x faster$0.186 vs $0.251
6-panel comic page2.3 min2 minNVIDIA B300 1.16x faster$0.233 vs $0.234
Character sheet, 12 poses2.9 min2.2 minNVIDIA B300 1.28x faster$0.285 vs $0.257
10 short social clips5.6 min4.7 minNVIDIA B300 1.18x faster$0.558 vs $0.548
Full codebase review6.6 min6.3 minNVIDIA B300 1.05x faster$0.660 vs $0.730
40-product photo shoot8.4 min6.1 minNVIDIA B300 1.38x faster$0.836 vs $0.701
20 long-form articles10.5 min9.7 minNVIDIA B300 1.08x faster$1.044 vs $1.125
40-product shoot, start to finish12.1 min6.9 minNVIDIA B300 1.75x faster$1.201 vs $0.797
200-product catalogue cutout17.2 min3.7 minNVIDIA B300 4.62x faster$1.712 vs $0.430
100-photo restoration batch17.8 min10.2 minNVIDIA B300 1.75x faster$1.773 vs $1.178
100-photo restore and enlarge26.4 min11.9 minNVIDIA B300 2.21x faster$2.628 vs $1.378

Renting by the hour, NVIDIA B300 finishes 8 of 14 cheaper. The quicker card is not automatically the cheaper way to get the work done.

Cost to rent

CardPer hour
NVIDIA B200$5.980
NVIDIA B300$6.940

NVIDIA B200 is 1.16x cheaper per hour. Cheapest on-demand rate we see across RunPod and Vast.

Specifications compared

NVIDIA B200NVIDIA B300
VRAM192GB288GB
ArchitectureBlackwellBlackwell Ultra
Memory bandwidth8 TB/s8 TB/s
TDP1000 W1400 W
Launch MSRP$40,000$40,000
Release2025-02-012025-11-01

FAQ

Which is better for ai & machine learning: NVIDIA B200 or NVIDIA B300?
NVIDIA B300 performs better for ai & machine learning, winning 135 of 142 benchmarks in our suite with an average 12.4% advantage.
What are the main hardware differences between NVIDIA B200 and NVIDIA B300?
NVIDIA B200 has 192GB VRAM and a 1000W TDP, while NVIDIA B300 has 288GB VRAM and a 1400W TDP.
Where is the biggest performance difference between NVIDIA B200 and NVIDIA B300?
Swin2SR 4x Upscaler: NVIDIA B300 leads by roughly 387% (11.97 vs 58.29 images/min) in our testing.

NVIDIA B200 full review · NVIDIA B300 full review · All AI & Machine Learning rankings