NVIDIA H200, AI & Machine Learning Benchmarks & Specs

141GB · AI Score 65.0/100 · first-party measured on 12 AI workloads

65 AI Score ✓ Measured

Every number on this page is first-party: NVIDIA H200 was run on our pinned 12-workload AI suite on 2026-07-10, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) NVIDIA H200 delivers about 268.31 tokens/sec. Stepping up to Qwen3 32B it holds roughly 76.58 tok/s. The full Llama 3.3 70B still runs, at about 42.66 tok/s. For image generation, SDXL runs at 18.58 it/s, and FLUX.1-dev at 4.44 it/s. All 12 workloads fit in 141GB. There is no model in our suite this card has to turn down. NVIDIA H200 isn't a retail purchase for most people. It's rented by the hour. You can run this exact card on RunPod.

AI & Machine Learning benchmark results

Text Generation tok/s 123

gemma-3-270m967.35
LFM2.5-1.2B946.53
Qwen2.5-Coder-0.5B920.36
Qwen2.5-0.5B914.58
SmolLM2-135M905.38
Llama 3.2 1B875.53
Qwen1.5-0.5B855.22
Qwen3-0.6B724.67
Qwen3 0.6B718.67
LFM2.5-8B-A1B583.96
Qwen3-1.7B568.85
Qwen3 1.7B563.52
WorkloadResultTelemetryData
gemma-3-270m967.35 tok/s
130 W35°CQ4_K_M
✓ Measured
LFM2.5-1.2B946.53 tok/s
131 W41°CQ4_K_M
✓ Measured
Qwen2.5-Coder-0.5B920.36 tok/s
132 W37°CQ4_K_M
✓ Measured
Qwen2.5-0.5B914.58 tok/s
135 W36°CQ4_K_M
✓ Measured
SmolLM2-135M905.38 tok/s
129 W39°CQ4_K_M
✓ Measured
Llama 3.2 1B875.53 tok/s
144 W37°CQ4_K_M
✓ Measured
Qwen1.5-0.5B855.22 tok/s
128 W37°CQ4_K_M
✓ Measured
Qwen3-0.6B724.67 tok/s
136 W37°CQ4_K_M
✓ Measured
Qwen3 0.6B718.67 tok/s
103 W36°CQ4_K_M
✓ Measured
LFM2.5-8B-A1B583.96 tok/s
148 W39°CQ4_K_M
✓ Measured
Qwen3-1.7B568.85 tok/s
165 W43°CQ4_K_M
✓ Measured
Qwen3 1.7B563.52 tok/s
100 W37°CQ4_K_M
✓ Measured
Qwen2.5-1.5B549.36 tok/s
163 W37°CQ4_K_M
✓ Measured
Qwen2-1.5B546.47 tok/s
154 W40°CQ4_K_M
✓ Measured
Qwen2.5-Coder-1.5B545.54 tok/s
158 W41°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 1.5B542.95 tok/s
108 W38°CQ4_K_M
✓ Measured
gemma-3-1b513.58 tok/s
153 W36°CQ4_K_M
✓ Measured
Hermes-3-Llama-3.2-3B430.79 tok/s
173 W42°CQ4_K_M
✓ Measured
Llama-3.2-3B-Instruct-uncensored429.78 tok/s
186 W45°CQ4_K_M
✓ Measured
Llama 3.2 3B424.61 tok/s
173 W39°CQ4_K_M
✓ Measured
SmolLM3-3B404.27 tok/s
187 W43°CQ4_K_M
✓ Measured
SmolLM3 3B402.24 tok/s
131 W40°CQ4_K_M
✓ Measured
Qwen2.5-Coder-3B401.11 tok/s
179 W43°CQ4_K_M
✓ Measured
Qwen2.5-3B400.48 tok/s
169 W39°CQ4_K_M
✓ Measured
gemma-2-2b398.48 tok/s
169 W39°CQ4_K_M
✓ Measured
gemma-2-2b-it-abliterated397.43 tok/s
171 W41°CQ4_K_M
✓ Measured
Phi-4-mini395.3 tok/s
205 W44°CQ4_K_M
✓ Measured
Phi-4 Mini 3.8B392.38 tok/s
193 W42°CQ4_K_M
✓ Measured
AI21-Jamba-Reasoning-3B370.66 tok/s
180 W41°CQ4_K_M
✓ Measured
gpt-oss-20b354.23 tok/s
162 W38°CQ4_K_M
✓ Measured
Phi-3.5-mini348.14 tok/s
193 W39°CQ4_K_M
✓ Measured
phi-2347.08 tok/s
191 W40°CQ4_K_M
✓ Measured
Nemotron-3-Nano-30B-A3B327.48 tok/s
156 W42°CQ4_K_M
✓ Measured
Qwen3-4B-Thinking-2507318.94 tok/s
219 W42°CQ4_K_M
✓ Measured
Qwen3 4B318.84 tok/s
3.3 GB peak107 W43°C2.98 tok/WQ4_K_M
✓ Measured
Qwen3-4B-Instruct-2507318.67 tok/s
210 W43°CQ4_K_M
✓ Measured
Qwen3-4B-Instruct-2507318.27 tok/s
199 W43°CQ4_K_M
✓ Measured
DeepSeek-Coder-V2-Lite312.1 tok/s
166 W40°CQ4_K_M
✓ Measured
Qwen3-Coder 30B A3B297.9 tok/s
124 W38°CQ4_K_M
✓ Measured
Qwen3-30B-A3B293.47 tok/s
158 W40°CQ4_K_M
✓ Measured
Qwen3 30B A3B292.11 tok/s
143 W43°CQ4_K_M
✓ Measured
Gemma 3 4B291.76 tok/s
188 W40°CQ4_K_M
✓ Measured
Llama-2-7B290.81 tok/s
218 W45°CQ4_K_M
✓ Measured
Mistral-7B-Instruct-v0.1282.64 tok/s
231 W45°CQ4_K_M
✓ Measured
Mistral-7B-Instruct-v0.3282.39 tok/s
233 W44°CQ4_K_M
✓ Measured
Mistral-7B-Instruct-v0.2282.32 tok/s
248 W44°CQ4_K_M
✓ Measured
Mistral 7B v0.3278.5 tok/s
105 W43°CQ4_K_M
✓ Measured
Qwen2.5-Coder-7B-Instruct-abliterated270.92 tok/s
239 W44°CQ4_K_M
✓ Measured
Dolphin X1 8B268.36 tok/s
230 W43°CQ4_K_M
✓ Measured
Llama 3.1 8B268.31 tok/s
5 GB peak213 W49°C1.26 tok/WQ4_K_M
✓ Measured
dolphin-2.9-llama3-8b267.96 tok/s
241 W45°CQ4_K_M
✓ Measured
L3-8B-Stheno-v3.2267.92 tok/s
230 W45°CQ4_K_M
✓ Measured
Dolphin 3.0 Llama 3.1 8B267.84 tok/s
224 W42°CQ4_K_M
✓ Measured
DarkIdol-Llama-3.1-8B-Instruct-1.2-Uncensored267.82 tok/s
219 W48°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 7B267.1 tok/s
110 W44°CQ4_K_M
✓ Measured
Qwen2.5-Coder 7B266.36 tok/s
132 W43°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill Llama 8B265.25 tok/s
131 W45°CQ4_K_M
✓ Measured
Qwen2.5-7B265.15 tok/s
233 W42°CQ4_K_M
✓ Measured
Meta-Llama-3.1-8B262.93 tok/s
231 W42°CQ4_K_M
✓ Measured
Dolphin X1 Trinity Nano 6B259.81 tok/s
144 W34°CQ4_K_M
✓ Measured
Qwen3-8B251.04 tok/s
222 W47°CQ4_K_M
✓ Measured
Josiefied-Qwen3-8B-abliterated-v1250.88 tok/s
249 W44°CQ4_K_M
✓ Measured
DeepSeek-R1-0528-Qwen3-8B250.61 tok/s
226 W42°CQ4_K_M
✓ Measured
Qwen3 8B247.98 tok/s
127 W45°CQ4_K_M
✓ Measured
Qwen3 30B A3B (Q3_K_M)246.35 tok/s
176 W42°CQ3_K_M
✓ Measured
KAT-Coder-V2.5-Dev236.31 tok/s
163 W39°CQ4_K_M
✓ Measured
Qwen-AgentWorld-35B-A3B225.31 tok/s
164 W37°CQ4_K_M
✓ Measured
Ornith-1.0-9B222.94 tok/s
228 W43°CQ4_K_M
✓ Measured
Ornith-1.0-35B208.51 tok/s
160 W37°CQ4_K_M
✓ Measured
Qwen3-Next-80B-A3B-Thinking201.12 tok/s
160 W40°CQ4_K_M
✓ Measured
Qwen3-Coder-Next-abliterated197.14 tok/s
153 W39°CQ4_K_M
✓ Measured
Qwen3-Coder-Next196.65 tok/s
155 W39°CQ4_K_M
✓ Measured
Qwen3-Coder-Next195.21 tok/s
151 W39°CQ4_K_M
✓ Measured
Qwen3-Next-80B-A3B-Thinking194.82 tok/s
148 W39°CQ4_K_M
✓ Measured
Qwen3-Next-80B-A3B192.52 tok/s
149 W40°CQ4_K_M
✓ Measured
GLM-4.7-Flash188.57 tok/s
174 W39°CQ4_K_M
✓ Measured
NemoMix-Unleashed-12B181.53 tok/s
279 W46°CQ4_K_M
✓ Measured
Mistral-Nemo-Instruct-2407181.3 tok/s
264 W45°CQ4_K_M
✓ Measured
gemma-2-9b180.47 tok/s
263 W45°CQ4_K_M
✓ Measured
GLM-4.7-Flash-REAP-23B-A3B171.45 tok/s
175 W40°CQ4_K_M
✓ Measured
Phi-4 14B168.2 tok/s
203 W49°CQ4_K_M
✓ Measured
Qwen3-14B156.3 tok/s
285 W46°CQ4_K_M
✓ Measured
Hermes-4-14B156.26 tok/s
299 W47°CQ4_K_M
✓ Measured
Qwen3 14B154.51 tok/s
123 W48°CQ4_K_M
✓ Measured
Gemma 3 12B153.26 tok/s
151 W47°CQ4_K_M
✓ Measured
Gemma 4 12B151.05 tok/s
100 W45°CQ4_K_M
✓ Measured
Qwen2.5-14B148.76 tok/s
286 W45°CQ4_K_M
✓ Measured
EVA-Qwen2.5-14B-v0.2148.74 tok/s
275 W46°CQ4_K_M
✓ Measured
Uncensored148.7 tok/s
289 W47°CQ4_K_M
✓ Measured
Qwen2.5-Coder-14B-Instruct-abliterated148.48 tok/s
301 W46°CQ4_K_M
✓ Measured
Qwen2.5-Coder 14B148.43 tok/s
9 GB peak123 W49°C1.2 tok/WQ4_K_M
✓ Measured
DeepSeek-R1 Distill 14B146.82 tok/s
203 W48°CQ4_K_M
✓ Measured
Phi-4 14B (Q3_K_M)139.97 tok/s
322 W49°CQ3_K_M
✓ Measured
StarCoder2 15B136.29 tok/s
214 W44°CQ4_K_M
✓ Measured
Gemma 3 12B (Q3_K_M)127.04 tok/s
297 W47°CQ3_K_M
✓ Measured
DeepSeek-R1 Distill 14B (Q3_K_M)119.66 tok/s
314 W48°CQ3_K_M
✓ Measured
Codestral 22B111.34 tok/s
171 W46°CQ4_K_M
✓ Measured
Dolphin-Mistral-24B-Venice-Edition109.2 tok/s
323 W48°CQ4_K_M
✓ Measured
Cydonia-24B-v4.3109.2 tok/s
305 W49°CQ4_K_M
✓ Measured
Dolphin Mistral 24B Venice109.14 tok/s
110 W42°CQ4_K_M
✓ Measured
Dolphin 3.0 R1 Mistral 24B109.06 tok/s
257 W47°CQ4_K_M
✓ Measured
Devstral Small 24B109.05 tok/s
138 W46°CQ4_K_M
✓ Measured
Mistral Small 24B107.84 tok/s
208 W51°CQ4_K_M
✓ Measured
Codestral 22B (Q3_K_M)88.03 tok/s
360 W51°CQ3_K_M
✓ Measured
Mistral Small 24B (Q3_K_M)85.1 tok/s
359 W50°CQ3_K_M
✓ Measured
Gemma 3 27B84.75 tok/s
227 W43°CQ4_K_M
✓ Measured
Olmo-3.1-32B-Think78.46 tok/s
336 W49°CQ4_K_M
✓ Measured
Dolphin 2.9.1 Yi 1.5 34B76.7 tok/s
221 W43°CQ4_K_M
✓ Measured
Qwen3 32B76.58 tok/s
19 GB peak119 W52°C0.65 tok/WQ4_K_M
✓ Measured
DeepSeek-R1 Distill 32B75.8 tok/s
244 W46°CQ4_K_M
✓ Measured
Qwen2.5-Coder 32B75.77 tok/s
260 W44°CQ4_K_M
✓ Measured
QwQ 32B75.77 tok/s
225 W43°CQ4_K_M
✓ Measured
Qwen2.5-32B75.76 tok/s
336 W49°CQ4_K_M
✓ Measured
DeepSeek-R1-Distill-Qwen-32B-abliterated75.71 tok/s
336 W50°CQ4_K_M
✓ Measured
Qwen2.5-Coder 32B (Q3_K_M)58.87 tok/s
370 W51°CQ3_K_M
✓ Measured
Qwen2.5-72B43.4 tok/s
365 W54°CQ4_K_M
✓ Measured
Hermes-4-70B42.77 tok/s
391 W55°CQ4_K_M
✓ Measured
DeepSeek-R1-Distill-Llama-70B42.76 tok/s
375 W54°CQ4_K_M
✓ Measured
Llama-3.3-70B-Instruct-abliterated42.75 tok/s
365 W53°CQ4_K_M
✓ Measured
Meta-Llama-3.1-70B42.72 tok/s
370 W55°CQ4_K_M
✓ Measured
Llama 3.3 70B42.66 tok/s
40.1 GB peak218 W58°C0.2 tok/WQ4_K_M
✓ Measured
Laguna-XS-2.1✕ Won't fit needs ~26 GBVRAM-gated at this precision✓ Measured
Nanbeige4.2-3B✕ Won't fit needs ~4 GBVRAM-gated at this precision✓ Measured

Image Generation images/min 6

FLUX.1 Schnell60.51
Z-Image Turbo (1024px)40.74
Stable Diffusion XL37.16
Z-Image Turbo23.175
FLUX.1 dev9.514
Krea 2 Turbo7.25
WorkloadResultTelemetryData
FLUX.1 Schnell60.51 images/min
567 W52°C
✓ Measured
Z-Image Turbo (1024px)40.74 images/min
667 W53°C
✓ Measured
Stable Diffusion XL37.16 images/min
16.2 GB peak604 W58°C1.6 s/img
✓ Measured
Z-Image Turbo23.18 images/min
26 GB peak665 W61°C2.6 s/img
✓ Measured
FLUX.1 dev9.51 images/min
36.9 GB peak687 W63°C6.3 s/img
✓ Measured
Krea 2 Turbo7.25 images/min
659 W65°C
✓ Measured

Fine-Tuning train tok/s 4

SmolLM2 1.7B LoRA17650.4
TinyLlama 1.1B LoRA16222.4
Qwen2.5 1.5B LoRA14931.7
Qwen2.5 7B LoRA8855.4
WorkloadResultTelemetryData
SmolLM2 1.7B LoRA17650.4 train tok/s
357 W51°C
✓ Measured
TinyLlama 1.1B LoRA16222.4 train tok/s
261 W46°C
✓ Measured
Qwen2.5 1.5B LoRA14931.7 train tok/s
308 W48°C
✓ Measured
Qwen2.5 7B LoRA8855.4 train tok/s
592 W60°C
✓ Measured

LLM Serving serve tok/s 4

TinyLlama 1.1B served9137.7
Qwen2.5 1.5B served7541.6
SmolLM2 1.7B served7107.8
Qwen2.5 7B served4394.8
WorkloadResultTelemetryData
TinyLlama 1.1B served9137.7 serve tok/s
203 W37°C
✓ Measured
Qwen2.5 1.5B served7541.6 serve tok/s
219 W38°C
✓ Measured
SmolLM2 1.7B served7107.8 serve tok/s
237 W39°C
✓ Measured
Qwen2.5 7B served4394.8 serve tok/s
467 W44°C
✓ Measured

Image Editing images/min 2

WorkloadResultTelemetryData
FLUX.1 Kontext dev4.39 images/min
35.7 GB peak689 W69°C13.7 s/img
✓ Measured
Qwen-Image-Edit3.76 images/min
60.4 GB peak685 W68°C15.9 s/img
✓ Measured

Video Generation frames/s 2

WorkloadResultTelemetryData
LTX-Video (distilled)17.87 frames/s
60.3 GB peak601 W59°C5.4 s/clip
✓ Measured
Wan 2.2 5B (720p)1.41 frames/s
37.9 GB peak685 W68°C34.9 s/clip
✓ Measured

Image to 3D assets/hour 2

WorkloadResultTelemetryData
TRELLIS Image-to-3D818.8 assets/hour
317 W56°C
✓ Measured
TRELLIS.2 Image-to-3D (1536³ max quality)65.8 assets/hour
643 W67°C
✓ Measured

Depth Estimation images/min 2

WorkloadResultTelemetryData
Depth Anything V2 Small1198.58 images/min
125 W32°C
✓ Measured
Depth Anything V2 Large1040.81 images/min
124 W32°C
✓ Measured

Segmentation images/min 2

WorkloadResultTelemetryData
SAM ViT-Base1469.34 images/min
141 W33°C
✓ Measured
SAM ViT-Huge325.19 images/min
300 W43°C
✓ Measured

Background Removal images/min 1

WorkloadResultTelemetryData
BiRefNet1457.73 images/min
134 W31°C
✓ Measured

Upscaling images/min 1

WorkloadResultTelemetryData
Swin2SR 4x Upscaler25.85 images/min
214 W37°C
✓ Measured

Text to Speech x realtime 1

WorkloadResultTelemetryData
Kokoro TTS 82M170.82 x realtime
120 W31°C
✓ Measured

Music Generation x realtime 1

WorkloadResultTelemetryData
MusicGen Small1.87 x realtime
146 W32°C
✓ Measured

Speech to Text x realtime 1

WorkloadResultTelemetryData
Whisper large-v3163.78 x realtime
159 W35°C
✓ Measured
How we measured this. Every result comes from our own pinned, reproducible AI suite, 12 workloads: the Qwen3-4B to Llama-70B LLM ladder (llama.cpp, Q4_K_M), SDXL / Z-Image / FLUX-dev generation, FLUX-Kontext / Qwen-Edit editing, and LTX / Wan video, run first-party on rented hardware with under 0.5% run-to-run variance. Peak VRAM, power draw, temperature and tokens-per-watt are captured per workload. “Won’t fit” rows are real data: where a model exceeds the card’s VRAM at the tested precision we record a hard gate rather than silently dropping to a smaller quant. Measured 2026-07-10 · harness 2.0.0.

NVIDIA H200 specifications

ArchitectureHopper
CUDA cores16,896
VRAM141GB HBM3e
Memory bus6144-bit
Memory bandwidth4800 GB/s
Boost clock1,980 MHz
TDP700 W
ProcessTSMC 4N
InterfaceSXM5
Release date2024-03-18
Launch MSRP$31,000

Verdict, nothing in our suite slows it down

NVIDIA H200 scores 65.0/100, #6 of 102. It ran all 12 workloads. Every figure here is our own measurement.

Relative performance: where the NVIDIA H200 lands

100% = this card, AI & Machine Learning headline metric (AI Score). #6 of 21 desktop cards in this vertical.

GPURelative%AI Score
NVIDIA B200
120%78
NVIDIA B100
103%67
NVIDIA H100 NVL
103%67
NVIDIA GH200 Grace Hopper
101%65.6
NVIDIA H200
100%65
NVIDIA H100 80GB HBM3
96%62.6
NVIDIA H800 80GB
96%62.6
NVIDIA RTX PRO 6000 Blackwell Server Edition
77%49.9
NVIDIA H100 PCIe
71%46.4

← All AI & Machine Learning GPU rankings

What this card can build

Whole-job timings, composed from our measured per-model results on this card.

WorkflowTimeEnergyBasis
50-image depth pass6 s0.1 Whmeasured
30-minute podcast pass46 s0.85 Whall 3 stages measured
500-image masking run1.7 min7.7 Whmeasured
20-asset 3D game kit2.4 min13.15 Whall 2 stages measured
6-panel comic page3.2 min23.21 Whall 3 stages measured
60-second AI short film3.9 min18.03 Whall 3 stages measured
Character sheet, 12 poses3.9 min32.58 Whall 2 stages measured
24-frame storyboard3.9 min12.08 Whall 2 stages measured
60-second AI short film, narrated4.1 min18.04 Whall 4 stages measured
Full codebase review6.7 min13.86 Whmeasured
200-product catalogue cutout8 min27.94 Whall 2 stages measured
10 short social clips10 min71.1 Whall 3 stages measured
20 long-form articles10.9 min39.82 Whmeasured
40-product photo shoot11.2 min115.44 Whall 2 stages measured
40-product shoot, start to finish12.9 min121.03 Whall 4 stages measured
100-photo restoration batch23.3 min261.51 Whmeasured
100-photo restore and enlarge27.2 min275.33 Whall 2 stages measured

Rent or buy?

This card is $31,000 to buy. The cheapest listed rate on RunPod is $3.590/hour, but that is the floor: we budget $4.308/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 7,196 GPU-hours. Below it you are paying for idle silicon.

How you would use itGPU-hours a yearRental cost a yearTime to break even
2 hours a day, hobby730$3,1459.9 years
8 hours a day, working on it2,920$12,5792.5 years
24/7, always-on agent8,760$37,7389.9 months

At hobby usage this card is very unlikely to pay for itself before it is superseded. Rent it. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.

Rental price

$3.590/hr+0.0% since 2026-08-14low $3.590 · high $3.590

Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.