NVIDIA B300 vs NVIDIA H200, AI & Machine Learning Comparison

NVIDIA B300
NVIDIA B300
vs
NVIDIA H200
NVIDIA H200

NVIDIA B300 wins 125 of 142 benchmarks, averaging 12.9% faster.

Both cards were measured first-party on our bench, same suite, same test rig.

What the numbers say

The gap is widest in FLUX.1 Kontext dev, where NVIDIA B300 leads by 135% (10.329 vs 4.393 images/min); the closest fight is Qwen3 0.6B (0% apart).

Benchmark results head-to-head

BenchmarkNVIDIA B300NVIDIA H200Difference
Qwen3 4B tok/s333.34318.84+5%
Llama 3.1 8B tok/s287.23268.31+7%
Qwen2.5-Coder 14B tok/s158.49148.43+7%
Qwen3 32B tok/s83.6876.58+9%
Llama 3.3 70B tok/s47.9742.66+12%
Stable Diffusion XL images/min29.237.16-21%
Z-Image Turbo images/min39.67523.175+71%
FLUX.1 dev images/min20.8079.514+119%
FLUX.1 Kontext dev images/min10.3294.393+135%
Qwen-Image-Edit images/min8.143.76+116%
LTX-Video (distilled) frames/s31.8717.87+78%
Wan 2.2 5B (720p) frames/s2.941.41+109%
DeepSeek-R1 Distill Llama 8B tok/s283.77265.25+7%
DeepSeek-R1 Distill 1.5B tok/s562.14542.95+4%
DeepSeek-R1 Distill 14B tok/s156.33146.82+6%
DeepSeek-R1 Distill 7B tok/s284.77267.1+7%
Gemma 3 12B tok/s161.27153.26+5%
Gemma 3 4B tok/s301.07291.76+3%
Gemma 4 12B tok/s155.96151.05+3%
Llama 3.2 1B tok/s889.19875.53+2%
Llama 3.2 3B tok/s432.24424.61+2%
Mistral 7B v0.3 tok/s298.95278.5+7%
Mistral Small 24B tok/s119.3107.84+11%
Phi-4 14B tok/s177.41168.2+5%
Phi-4 Mini 3.8B tok/s398.16392.38+1%
Qwen2.5-Coder 7B tok/s286.5266.36+8%
Qwen3 0.6B tok/s718.58718.67-0%
Qwen3 1.7B tok/s555.1563.52-1%
Qwen3 14B tok/s165.44154.51+7%
Qwen3 30B A3B tok/s284.66292.11-3%
Qwen3 8B tok/s261.4247.98+5%
SmolLM3 3B tok/s411.25402.24+2%
Codestral 22B tok/s119.81111.34+8%
DeepSeek-R1 Distill 32B tok/s82.975.8+9%
Devstral Small 24B tok/s121.18109.05+11%
Dolphin 2.9.1 Yi 1.5 34B tok/s85.1576.7+11%
Dolphin Mistral 24B Venice tok/s120.93109.14+11%
Dolphin X1 8B tok/s287.02268.36+7%
Dolphin 3.0 Llama 3.1 8B tok/s285.88267.84+7%
Dolphin 3.0 R1 Mistral 24B tok/s120.95109.06+11%
Gemma 3 27B tok/s93.4484.75+10%
Qwen2.5-Coder 32B tok/s82.975.77+9%
Qwen3-Coder 30B A3B tok/s284.88297.9-4%
QwQ 32B tok/s82.9175.77+9%
StarCoder2 15B tok/s144.05136.29+6%
FLUX.1 Schnell images/min59.5460.51-2%
Z-Image Turbo (1024px) images/min48.6340.74+19%
BiRefNet images/min1064.731457.73-27%
Depth Anything V2 Large images/min1069.951040.81+3%
Depth Anything V2 Small images/min1491.61198.58+24%
SAM ViT-Base images/min1719.591469.34+17%
SAM ViT-Huge images/min399.37325.19+23%
Swin2SR 4x Upscaler images/min58.2925.85+125%
Qwen2.5 1.5B LoRA train tok/s22351.114931.7+50%
Qwen2.5 7B LoRA train tok/s14323.18855.4+62%
SmolLM2 1.7B LoRA train tok/s26114.317650.4+48%
TinyLlama 1.1B LoRA train tok/s26938.216222.4+66%
AI21-Jamba-Reasoning-3B tok/s380.88370.66+3%
Olmo-3.1-32B-Think tok/s87.8578.46+12%
Codestral 22B (Q3_K_M) tok/s107.9888.03+23%
Dolphin-Mistral-24B-Venice-Edition tok/s121.87109.2+12%
DarkIdol-Llama-3.1-8B-Instruct-1.2-Uncensored tok/s292.55267.82+9%
DeepSeek-Coder-V2-Lite tok/s340.53312.1+9%
DeepSeek-R1-0528-Qwen3-8B tok/s270.35250.61+8%
DeepSeek-R1-Distill-Llama-70B tok/s48.1142.76+13%
DeepSeek-R1 Distill 14B (Q3_K_M) tok/s144.74119.66+21%
DeepSeek-R1-Distill-Qwen-32B-abliterated tok/s83.1975.71+10%
dolphin-2.9-llama3-8b tok/s290.47267.96+8%
Dolphin X1 Trinity Nano 6B tok/s257.93259.81-1%
EVA-Qwen2.5-14B-v0.2 tok/s159.46148.74+7%
gemma-2-2b-it-abliterated tok/s418.18397.43+5%
gemma-2-2b tok/s417.84398.48+5%
gemma-2-9b tok/s202.22180.47+12%
Gemma 3 12B (Q3_K_M) tok/s149.61127.04+18%
gemma-3-1b tok/s533.32513.58+4%
gemma-3-270m tok/s926.67967.35-4%
GLM-4.7-Flash-REAP-23B-A3B tok/s178.49171.45+4%
GLM-4.7-Flash tok/s196.52188.57+4%
Josiefied-Qwen3-8B-abliterated-v1 tok/s270.09250.88+8%
gpt-oss-20b tok/s355.56354.230%
Hermes-3-Llama-3.2-3B tok/s446.78430.79+4%
Hermes-4-70B tok/s48.142.77+12%
SmolLM3-3B tok/s420.3404.27+4%
Qwen3-Coder-Next-abliterated tok/s199.12197.14+1%
KAT-Coder-V2.5-Dev tok/s239.66236.31+1%
L3-8B-Stheno-v3.2 tok/s293267.92+9%
Laguna-XS-2.1 tok/s00n/a
LFM2.5-1.2B tok/s955.13946.53+1%
LFM2.5-8B-A1B tok/s596.16583.96+2%
Llama-2-7B tok/s315.9290.81+9%
Llama-3.2-3B-Instruct-uncensored tok/s448.94429.78+4%
Llama-3.3-70B-Instruct-abliterated tok/s48.142.75+13%
Meta-Llama-3.1-70B tok/s48.1242.72+13%
Meta-Llama-3.1-8B tok/s293262.93+11%
Phi-4-mini tok/s409.58395.3+4%
Mistral-7B-Instruct-v0.1 tok/s306.44282.64+8%
Mistral-7B-Instruct-v0.2 tok/s306.79282.32+9%
Mistral-7B-Instruct-v0.3 tok/s306.46282.39+9%
Mistral-Nemo-Instruct-2407 tok/s196.92181.3+9%
Mistral Small 24B (Q3_K_M) tok/s108.3785.1+27%
Nanbeige4.2-3B tok/s00n/a
NemoMix-Unleashed-12B tok/s197.06181.53+9%
Nemotron-3-Nano-30B-A3B tok/s345.31327.48+5%
Hermes-4-14B tok/s169.51156.26+8%
Ornith-1.0-35B tok/s232.36208.51+11%
Ornith-1.0-9B tok/s240.67222.94+8%
phi-2 tok/s356.65347.08+3%
Phi-3.5-mini tok/s352.25348.14+1%
Phi-4 14B (Q3_K_M) tok/s165.06139.97+18%
Qwen-AgentWorld-35B-A3B tok/s233.04225.31+3%
Qwen3-0.6B tok/s742.64724.67+2%
Qwen3-1.7B tok/s575.97568.85+1%
Qwen3-14B tok/s169.52156.3+8%
Qwen3-30B-A3B tok/s291.82293.47-1%
Qwen3-4B-Instruct-2507 tok/s339.78318.67+7%
Qwen3-8B tok/s270.35251.04+8%
Qwen3-Coder-Next tok/s196.34195.21+1%
Qwen3-Next-80B-A3B-Thinking tok/s194.98194.820%
Qwen1.5-0.5B tok/s894.84855.22+5%
Qwen2-1.5B tok/s583.79546.47+7%
Qwen2.5-0.5B tok/s843.4914.58-8%
Qwen2.5-1.5B tok/s584.96549.36+6%
Uncensored tok/s159.48148.7+7%
Qwen2.5-14B tok/s159.33148.76+7%
Qwen2.5-32B tok/s83.4175.76+10%
Qwen2.5-3B tok/s409.56400.48+2%
Qwen2.5-72B tok/s48.2843.4+11%
Qwen2.5-7B tok/s293.63265.15+11%
Qwen2.5-Coder-0.5B tok/s843.44920.36-8%
Qwen2.5-Coder-1.5B tok/s584.5545.54+7%
Qwen2.5-Coder-14B-Instruct-abliterated tok/s159.5148.48+7%
Qwen2.5-Coder 32B (Q3_K_M) tok/s74.2858.87+26%
Qwen2.5-Coder-3B tok/s410.06401.11+2%
Qwen2.5-Coder-7B-Instruct-abliterated tok/s292.21270.92+8%
Qwen3 30B A3B (Q3_K_M) tok/s268.71246.35+9%
Qwen3-4B-Instruct-2507 tok/s339.72318.27+7%
Qwen3-4B-Thinking-2507 tok/s339.74318.94+7%
Qwen3-Coder-Next tok/s199.39196.65+1%
Qwen3-Next-80B-A3B-Thinking tok/s200.55201.12-0%
Qwen3-Next-80B-A3B tok/s191.32192.52-1%
SmolLM2-135M tok/s855.87905.38-5%
Cydonia-24B-v4.3 tok/s121.9109.2+12%

Whole-job comparison

How long each card takes to finish a complete pipeline, not just one model. NVIDIA B300 is faster on 14 of 14; NVIDIA H200 on 0.

WorkflowNVIDIA B300NVIDIA H200DifferenceCost per run
50-image depth pass5 s6 sNVIDIA B300 1.33x faster$0.009 vs $0.006
500-image masking run78 s1.7 minNVIDIA B300 1.29x faster$0.150 vs $0.100
24-frame storyboard1.5 min3.9 minNVIDIA B300 2.60x faster$0.174 vs $0.234
6-panel comic page2 min3.2 minNVIDIA B300 1.59x faster$0.234 vs $0.192
60-second AI short film2.2 min3.9 minNVIDIA B300 1.79x faster$0.251 vs $0.233
Character sheet, 12 poses2.2 min3.9 minNVIDIA B300 1.75x faster$0.257 vs $0.234
200-product catalogue cutout3.7 min8 minNVIDIA B300 2.15x faster$0.430 vs $0.479
10 short social clips4.7 min10 minNVIDIA B300 2.11x faster$0.548 vs $0.598
40-product photo shoot6.1 min11.2 minNVIDIA B300 1.84x faster$0.701 vs $0.668
Full codebase review6.3 min6.7 minNVIDIA B300 1.07x faster$0.730 vs $0.403
40-product shoot, start to finish6.9 min12.9 minNVIDIA B300 1.87x faster$0.797 vs $0.770
20 long-form articles9.7 min10.9 minNVIDIA B300 1.12x faster$1.125 vs $0.655
100-photo restoration batch10.2 min23.3 minNVIDIA B300 2.29x faster$1.178 vs $1.396
100-photo restore and enlarge11.9 min27.2 minNVIDIA B300 2.29x faster$1.378 vs $1.629

Renting by the hour, NVIDIA H200 finishes 9 of 14 cheaper. The quicker card is not automatically the cheaper way to get the work done.

Cost to rent

CardPer hour
NVIDIA B300$6.940
NVIDIA H200$3.590

NVIDIA H200 is 1.93x cheaper per hour. Cheapest on-demand rate we see across RunPod and Vast.

Specifications compared

NVIDIA B300NVIDIA H200
VRAM288GB141GB
ArchitectureBlackwell UltraHopper
Memory bandwidth8 TB/s4800 GB/s
Boost clockn/a1,980 MHz
TDP1400 W700 W
Launch MSRP$40,000$31,000
Release2025-11-012024-03-18

FAQ

Which is better for ai & machine learning: NVIDIA B300 or NVIDIA H200?
NVIDIA B300 performs better for ai & machine learning, winning 125 of 142 benchmarks in our suite with an average 12.9% advantage.
What are the main hardware differences between NVIDIA B300 and NVIDIA H200?
NVIDIA B300 has 288GB VRAM and a 1400W TDP, while NVIDIA H200 has 141GB VRAM and a 700W TDP.
Where is the biggest performance difference between NVIDIA B300 and NVIDIA H200?
FLUX.1 Kontext dev: NVIDIA B300 leads by roughly 135% (10.329 vs 4.393 images/min) in our testing.

NVIDIA B300 full review · NVIDIA H200 full review · All AI & Machine Learning rankings