141GB · AI Score 65.0/100 · first-party measured on 12 AI workloads
65 AI Score ✓ Measured
Every number on this page is first-party: NVIDIA H200 was run on our pinned 12-workload AI suite on 2026-07-10, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) NVIDIA H200 delivers about 268.31 tokens/sec. Stepping up to Qwen3 32B it holds roughly 76.58 tok/s. The full Llama 3.3 70B still runs, at about 42.66 tok/s. For image generation, SDXL runs at 18.58 it/s, and FLUX.1-dev at 4.44 it/s. All 12 workloads fit in 141GB. There is no model in our suite this card has to turn down. NVIDIA H200 isn't a retail purchase for most people. It's rented by the hour. You can run this exact card on RunPod.
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| gemma-3-270m | 967.35 tok/s | 130 W35°CQ4_K_M | ✓ Measured |
| LFM2.5-1.2B | 946.53 tok/s | 131 W41°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder-0.5B | 920.36 tok/s | 132 W37°CQ4_K_M | ✓ Measured |
| Qwen2.5-0.5B | 914.58 tok/s | 135 W36°CQ4_K_M | ✓ Measured |
| SmolLM2-135M | 905.38 tok/s | 129 W39°CQ4_K_M | ✓ Measured |
| Llama 3.2 1B | 875.53 tok/s | 144 W37°CQ4_K_M | ✓ Measured |
| Qwen1.5-0.5B | 855.22 tok/s | 128 W37°CQ4_K_M | ✓ Measured |
| Qwen3-0.6B | 724.67 tok/s | 136 W37°CQ4_K_M | ✓ Measured |
| Qwen3 0.6B | 718.67 tok/s | 103 W36°CQ4_K_M | ✓ Measured |
| LFM2.5-8B-A1B | 583.96 tok/s | 148 W39°CQ4_K_M | ✓ Measured |
| Qwen3-1.7B | 568.85 tok/s | 165 W43°CQ4_K_M | ✓ Measured |
| Qwen3 1.7B | 563.52 tok/s | 100 W37°CQ4_K_M | ✓ Measured |
| Qwen2.5-1.5B | 549.36 tok/s | 163 W37°CQ4_K_M | ✓ Measured |
| Qwen2-1.5B | 546.47 tok/s | 154 W40°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder-1.5B | 545.54 tok/s | 158 W41°CQ4_K_M | ✓ Measured |
| DeepSeek-R1 Distill 1.5B | 542.95 tok/s | 108 W38°CQ4_K_M | ✓ Measured |
| gemma-3-1b | 513.58 tok/s | 153 W36°CQ4_K_M | ✓ Measured |
| Hermes-3-Llama-3.2-3B | 430.79 tok/s | 173 W42°CQ4_K_M | ✓ Measured |
| Llama-3.2-3B-Instruct-uncensored | 429.78 tok/s | 186 W45°CQ4_K_M | ✓ Measured |
| Llama 3.2 3B | 424.61 tok/s | 173 W39°CQ4_K_M | ✓ Measured |
| SmolLM3-3B | 404.27 tok/s | 187 W43°CQ4_K_M | ✓ Measured |
| SmolLM3 3B | 402.24 tok/s | 131 W40°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder-3B | 401.11 tok/s | 179 W43°CQ4_K_M | ✓ Measured |
| Qwen2.5-3B | 400.48 tok/s | 169 W39°CQ4_K_M | ✓ Measured |
| gemma-2-2b | 398.48 tok/s | 169 W39°CQ4_K_M | ✓ Measured |
| gemma-2-2b-it-abliterated | 397.43 tok/s | 171 W41°CQ4_K_M | ✓ Measured |
| Phi-4-mini | 395.3 tok/s | 205 W44°CQ4_K_M | ✓ Measured |
| Phi-4 Mini 3.8B | 392.38 tok/s | 193 W42°CQ4_K_M | ✓ Measured |
| AI21-Jamba-Reasoning-3B | 370.66 tok/s | 180 W41°CQ4_K_M | ✓ Measured |
| gpt-oss-20b | 354.23 tok/s | 162 W38°CQ4_K_M | ✓ Measured |
| Phi-3.5-mini | 348.14 tok/s | 193 W39°CQ4_K_M | ✓ Measured |
| phi-2 | 347.08 tok/s | 191 W40°CQ4_K_M | ✓ Measured |
| Nemotron-3-Nano-30B-A3B | 327.48 tok/s | 156 W42°CQ4_K_M | ✓ Measured |
| Qwen3-4B-Thinking-2507 | 318.94 tok/s | 219 W42°CQ4_K_M | ✓ Measured |
| Qwen3 4B | 318.84 tok/s | 3.3 GB peak107 W43°C2.98 tok/WQ4_K_M | ✓ Measured |
| Qwen3-4B-Instruct-2507 | 318.67 tok/s | 210 W43°CQ4_K_M | ✓ Measured |
| Qwen3-4B-Instruct-2507 | 318.27 tok/s | 199 W43°CQ4_K_M | ✓ Measured |
| DeepSeek-Coder-V2-Lite | 312.1 tok/s | 166 W40°CQ4_K_M | ✓ Measured |
| Qwen3-Coder 30B A3B | 297.9 tok/s | 124 W38°CQ4_K_M | ✓ Measured |
| Qwen3-30B-A3B | 293.47 tok/s | 158 W40°CQ4_K_M | ✓ Measured |
| Qwen3 30B A3B | 292.11 tok/s | 143 W43°CQ4_K_M | ✓ Measured |
| Gemma 3 4B | 291.76 tok/s | 188 W40°CQ4_K_M | ✓ Measured |
| Llama-2-7B | 290.81 tok/s | 218 W45°CQ4_K_M | ✓ Measured |
| Mistral-7B-Instruct-v0.1 | 282.64 tok/s | 231 W45°CQ4_K_M | ✓ Measured |
| Mistral-7B-Instruct-v0.3 | 282.39 tok/s | 233 W44°CQ4_K_M | ✓ Measured |
| Mistral-7B-Instruct-v0.2 | 282.32 tok/s | 248 W44°CQ4_K_M | ✓ Measured |
| Mistral 7B v0.3 | 278.5 tok/s | 105 W43°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder-7B-Instruct-abliterated | 270.92 tok/s | 239 W44°CQ4_K_M | ✓ Measured |
| Dolphin X1 8B | 268.36 tok/s | 230 W43°CQ4_K_M | ✓ Measured |
| Llama 3.1 8B | 268.31 tok/s | 5 GB peak213 W49°C1.26 tok/WQ4_K_M | ✓ Measured |
| dolphin-2.9-llama3-8b | 267.96 tok/s | 241 W45°CQ4_K_M | ✓ Measured |
| L3-8B-Stheno-v3.2 | 267.92 tok/s | 230 W45°CQ4_K_M | ✓ Measured |
| Dolphin 3.0 Llama 3.1 8B | 267.84 tok/s | 224 W42°CQ4_K_M | ✓ Measured |
| DarkIdol-Llama-3.1-8B-Instruct-1.2-Uncensored | 267.82 tok/s | 219 W48°CQ4_K_M | ✓ Measured |
| DeepSeek-R1 Distill 7B | 267.1 tok/s | 110 W44°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder 7B | 266.36 tok/s | 132 W43°CQ4_K_M | ✓ Measured |
| DeepSeek-R1 Distill Llama 8B | 265.25 tok/s | 131 W45°CQ4_K_M | ✓ Measured |
| Qwen2.5-7B | 265.15 tok/s | 233 W42°CQ4_K_M | ✓ Measured |
| Meta-Llama-3.1-8B | 262.93 tok/s | 231 W42°CQ4_K_M | ✓ Measured |
| Dolphin X1 Trinity Nano 6B | 259.81 tok/s | 144 W34°CQ4_K_M | ✓ Measured |
| Qwen3-8B | 251.04 tok/s | 222 W47°CQ4_K_M | ✓ Measured |
| Josiefied-Qwen3-8B-abliterated-v1 | 250.88 tok/s | 249 W44°CQ4_K_M | ✓ Measured |
| DeepSeek-R1-0528-Qwen3-8B | 250.61 tok/s | 226 W42°CQ4_K_M | ✓ Measured |
| Qwen3 8B | 247.98 tok/s | 127 W45°CQ4_K_M | ✓ Measured |
| Qwen3 30B A3B (Q3_K_M) | 246.35 tok/s | 176 W42°CQ3_K_M | ✓ Measured |
| KAT-Coder-V2.5-Dev | 236.31 tok/s | 163 W39°CQ4_K_M | ✓ Measured |
| Qwen-AgentWorld-35B-A3B | 225.31 tok/s | 164 W37°CQ4_K_M | ✓ Measured |
| Ornith-1.0-9B | 222.94 tok/s | 228 W43°CQ4_K_M | ✓ Measured |
| Ornith-1.0-35B | 208.51 tok/s | 160 W37°CQ4_K_M | ✓ Measured |
| Qwen3-Next-80B-A3B-Thinking | 201.12 tok/s | 160 W40°CQ4_K_M | ✓ Measured |
| Qwen3-Coder-Next-abliterated | 197.14 tok/s | 153 W39°CQ4_K_M | ✓ Measured |
| Qwen3-Coder-Next | 196.65 tok/s | 155 W39°CQ4_K_M | ✓ Measured |
| Qwen3-Coder-Next | 195.21 tok/s | 151 W39°CQ4_K_M | ✓ Measured |
| Qwen3-Next-80B-A3B-Thinking | 194.82 tok/s | 148 W39°CQ4_K_M | ✓ Measured |
| Qwen3-Next-80B-A3B | 192.52 tok/s | 149 W40°CQ4_K_M | ✓ Measured |
| GLM-4.7-Flash | 188.57 tok/s | 174 W39°CQ4_K_M | ✓ Measured |
| NemoMix-Unleashed-12B | 181.53 tok/s | 279 W46°CQ4_K_M | ✓ Measured |
| Mistral-Nemo-Instruct-2407 | 181.3 tok/s | 264 W45°CQ4_K_M | ✓ Measured |
| gemma-2-9b | 180.47 tok/s | 263 W45°CQ4_K_M | ✓ Measured |
| GLM-4.7-Flash-REAP-23B-A3B | 171.45 tok/s | 175 W40°CQ4_K_M | ✓ Measured |
| Phi-4 14B | 168.2 tok/s | 203 W49°CQ4_K_M | ✓ Measured |
| Qwen3-14B | 156.3 tok/s | 285 W46°CQ4_K_M | ✓ Measured |
| Hermes-4-14B | 156.26 tok/s | 299 W47°CQ4_K_M | ✓ Measured |
| Qwen3 14B | 154.51 tok/s | 123 W48°CQ4_K_M | ✓ Measured |
| Gemma 3 12B | 153.26 tok/s | 151 W47°CQ4_K_M | ✓ Measured |
| Gemma 4 12B | 151.05 tok/s | 100 W45°CQ4_K_M | ✓ Measured |
| Qwen2.5-14B | 148.76 tok/s | 286 W45°CQ4_K_M | ✓ Measured |
| EVA-Qwen2.5-14B-v0.2 | 148.74 tok/s | 275 W46°CQ4_K_M | ✓ Measured |
| Uncensored | 148.7 tok/s | 289 W47°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder-14B-Instruct-abliterated | 148.48 tok/s | 301 W46°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder 14B | 148.43 tok/s | 9 GB peak123 W49°C1.2 tok/WQ4_K_M | ✓ Measured |
| DeepSeek-R1 Distill 14B | 146.82 tok/s | 203 W48°CQ4_K_M | ✓ Measured |
| Phi-4 14B (Q3_K_M) | 139.97 tok/s | 322 W49°CQ3_K_M | ✓ Measured |
| StarCoder2 15B | 136.29 tok/s | 214 W44°CQ4_K_M | ✓ Measured |
| Gemma 3 12B (Q3_K_M) | 127.04 tok/s | 297 W47°CQ3_K_M | ✓ Measured |
| DeepSeek-R1 Distill 14B (Q3_K_M) | 119.66 tok/s | 314 W48°CQ3_K_M | ✓ Measured |
| Codestral 22B | 111.34 tok/s | 171 W46°CQ4_K_M | ✓ Measured |
| Dolphin-Mistral-24B-Venice-Edition | 109.2 tok/s | 323 W48°CQ4_K_M | ✓ Measured |
| Cydonia-24B-v4.3 | 109.2 tok/s | 305 W49°CQ4_K_M | ✓ Measured |
| Dolphin Mistral 24B Venice | 109.14 tok/s | 110 W42°CQ4_K_M | ✓ Measured |
| Dolphin 3.0 R1 Mistral 24B | 109.06 tok/s | 257 W47°CQ4_K_M | ✓ Measured |
| Devstral Small 24B | 109.05 tok/s | 138 W46°CQ4_K_M | ✓ Measured |
| Mistral Small 24B | 107.84 tok/s | 208 W51°CQ4_K_M | ✓ Measured |
| Codestral 22B (Q3_K_M) | 88.03 tok/s | 360 W51°CQ3_K_M | ✓ Measured |
| Mistral Small 24B (Q3_K_M) | 85.1 tok/s | 359 W50°CQ3_K_M | ✓ Measured |
| Gemma 3 27B | 84.75 tok/s | 227 W43°CQ4_K_M | ✓ Measured |
| Olmo-3.1-32B-Think | 78.46 tok/s | 336 W49°CQ4_K_M | ✓ Measured |
| Dolphin 2.9.1 Yi 1.5 34B | 76.7 tok/s | 221 W43°CQ4_K_M | ✓ Measured |
| Qwen3 32B | 76.58 tok/s | 19 GB peak119 W52°C0.65 tok/WQ4_K_M | ✓ Measured |
| DeepSeek-R1 Distill 32B | 75.8 tok/s | 244 W46°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder 32B | 75.77 tok/s | 260 W44°CQ4_K_M | ✓ Measured |
| QwQ 32B | 75.77 tok/s | 225 W43°CQ4_K_M | ✓ Measured |
| Qwen2.5-32B | 75.76 tok/s | 336 W49°CQ4_K_M | ✓ Measured |
| DeepSeek-R1-Distill-Qwen-32B-abliterated | 75.71 tok/s | 336 W50°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder 32B (Q3_K_M) | 58.87 tok/s | 370 W51°CQ3_K_M | ✓ Measured |
| Qwen2.5-72B | 43.4 tok/s | 365 W54°CQ4_K_M | ✓ Measured |
| Hermes-4-70B | 42.77 tok/s | 391 W55°CQ4_K_M | ✓ Measured |
| DeepSeek-R1-Distill-Llama-70B | 42.76 tok/s | 375 W54°CQ4_K_M | ✓ Measured |
| Llama-3.3-70B-Instruct-abliterated | 42.75 tok/s | 365 W53°CQ4_K_M | ✓ Measured |
| Meta-Llama-3.1-70B | 42.72 tok/s | 370 W55°CQ4_K_M | ✓ Measured |
| Llama 3.3 70B | 42.66 tok/s | 40.1 GB peak218 W58°C0.2 tok/WQ4_K_M | ✓ Measured |
| Laguna-XS-2.1 | ✕ Won't fit needs ~26 GB | VRAM-gated at this precision | ✓ Measured |
| Nanbeige4.2-3B | ✕ Won't fit needs ~4 GB | VRAM-gated at this precision | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| FLUX.1 Schnell | 60.51 images/min | 567 W52°C | ✓ Measured |
| Z-Image Turbo (1024px) | 40.74 images/min | 667 W53°C | ✓ Measured |
| Stable Diffusion XL | 37.16 images/min | 16.2 GB peak604 W58°C1.6 s/img | ✓ Measured |
| Z-Image Turbo | 23.18 images/min | 26 GB peak665 W61°C2.6 s/img | ✓ Measured |
| FLUX.1 dev | 9.51 images/min | 36.9 GB peak687 W63°C6.3 s/img | ✓ Measured |
| Krea 2 Turbo | 7.25 images/min | 659 W65°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| SmolLM2 1.7B LoRA | 17650.4 train tok/s | 357 W51°C | ✓ Measured |
| TinyLlama 1.1B LoRA | 16222.4 train tok/s | 261 W46°C | ✓ Measured |
| Qwen2.5 1.5B LoRA | 14931.7 train tok/s | 308 W48°C | ✓ Measured |
| Qwen2.5 7B LoRA | 8855.4 train tok/s | 592 W60°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| TinyLlama 1.1B served | 9137.7 serve tok/s | 203 W37°C | ✓ Measured |
| Qwen2.5 1.5B served | 7541.6 serve tok/s | 219 W38°C | ✓ Measured |
| SmolLM2 1.7B served | 7107.8 serve tok/s | 237 W39°C | ✓ Measured |
| Qwen2.5 7B served | 4394.8 serve tok/s | 467 W44°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| FLUX.1 Kontext dev | 4.39 images/min | 35.7 GB peak689 W69°C13.7 s/img | ✓ Measured |
| Qwen-Image-Edit | 3.76 images/min | 60.4 GB peak685 W68°C15.9 s/img | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| LTX-Video (distilled) | 17.87 frames/s | 60.3 GB peak601 W59°C5.4 s/clip | ✓ Measured |
| Wan 2.2 5B (720p) | 1.41 frames/s | 37.9 GB peak685 W68°C34.9 s/clip | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| TRELLIS Image-to-3D | 818.8 assets/hour | 317 W56°C | ✓ Measured |
| TRELLIS.2 Image-to-3D (1536³ max quality) | 65.8 assets/hour | 643 W67°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Depth Anything V2 Small | 1198.58 images/min | 125 W32°C | ✓ Measured |
| Depth Anything V2 Large | 1040.81 images/min | 124 W32°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| SAM ViT-Base | 1469.34 images/min | 141 W33°C | ✓ Measured |
| SAM ViT-Huge | 325.19 images/min | 300 W43°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| BiRefNet | 1457.73 images/min | 134 W31°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Swin2SR 4x Upscaler | 25.85 images/min | 214 W37°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Kokoro TTS 82M | 170.82 x realtime | 120 W31°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| MusicGen Small | 1.87 x realtime | 146 W32°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Whisper large-v3 | 163.78 x realtime | 159 W35°C | ✓ Measured |
| Architecture | Hopper |
| CUDA cores | 16,896 |
| VRAM | 141GB HBM3e |
| Memory bus | 6144-bit |
| Memory bandwidth | 4800 GB/s |
| Boost clock | 1,980 MHz |
| TDP | 700 W |
| Process | TSMC 4N |
| Interface | SXM5 |
| Release date | 2024-03-18 |
| Launch MSRP | $31,000 |
NVIDIA H200 scores 65.0/100, #6 of 102. It ran all 12 workloads. Every figure here is our own measurement.
100% = this card, AI & Machine Learning headline metric (AI Score). #6 of 21 desktop cards in this vertical.
| GPU | Relative | % | AI Score |
|---|---|---|---|
| NVIDIA B200 | 120% | 78 | |
| NVIDIA B100 | 103% | 67 | |
| NVIDIA H100 NVL | 103% | 67 | |
| NVIDIA GH200 Grace Hopper | 101% | 65.6 | |
| NVIDIA H200 | 100% | 65 | |
| NVIDIA H100 80GB HBM3 | 96% | 62.6 | |
| NVIDIA H800 80GB | 96% | 62.6 | |
| NVIDIA RTX PRO 6000 Blackwell Server Edition | 77% | 49.9 | |
| NVIDIA H100 PCIe | 71% | 46.4 |
← All AI & Machine Learning GPU rankings
Whole-job timings, composed from our measured per-model results on this card.
| Workflow | Time | Energy | Basis |
|---|---|---|---|
| 50-image depth pass | 6 s | 0.1 Wh | measured |
| 30-minute podcast pass | 46 s | 0.85 Wh | all 3 stages measured |
| 500-image masking run | 1.7 min | 7.7 Wh | measured |
| 20-asset 3D game kit | 2.4 min | 13.15 Wh | all 2 stages measured |
| 6-panel comic page | 3.2 min | 23.21 Wh | all 3 stages measured |
| 60-second AI short film | 3.9 min | 18.03 Wh | all 3 stages measured |
| Character sheet, 12 poses | 3.9 min | 32.58 Wh | all 2 stages measured |
| 24-frame storyboard | 3.9 min | 12.08 Wh | all 2 stages measured |
| 60-second AI short film, narrated | 4.1 min | 18.04 Wh | all 4 stages measured |
| Full codebase review | 6.7 min | 13.86 Wh | measured |
| 200-product catalogue cutout | 8 min | 27.94 Wh | all 2 stages measured |
| 10 short social clips | 10 min | 71.1 Wh | all 3 stages measured |
| 20 long-form articles | 10.9 min | 39.82 Wh | measured |
| 40-product photo shoot | 11.2 min | 115.44 Wh | all 2 stages measured |
| 40-product shoot, start to finish | 12.9 min | 121.03 Wh | all 4 stages measured |
| 100-photo restoration batch | 23.3 min | 261.51 Wh | measured |
| 100-photo restore and enlarge | 27.2 min | 275.33 Wh | all 2 stages measured |
This card is $31,000 to buy. The cheapest listed rate on RunPod is $3.590/hour, but that is the floor: we budget $4.308/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 7,196 GPU-hours. Below it you are paying for idle silicon.
| How you would use it | GPU-hours a year | Rental cost a year | Time to break even |
|---|---|---|---|
| 2 hours a day, hobby | 730 | $3,145 | 9.9 years |
| 8 hours a day, working on it | 2,920 | $12,579 | 2.5 years |
| 24/7, always-on agent | 8,760 | $37,738 | 9.9 months |
At hobby usage this card is very unlikely to pay for itself before it is superseded. Rent it. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.
Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.