80GB · AI Score 62.6/100 · first-party measured on 12 AI workloads
62.6 AI Score ✓ Measured
Every number on this page is first-party: NVIDIA H100 80GB HBM3 was run on our pinned 12-workload AI suite on 2026-07-10, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) NVIDIA H100 80GB HBM3 delivers about 261.83 tokens/sec. Stepping up to Qwen3 32B it holds roughly 74.07 tok/s. The full Llama 3.3 70B still runs, at about 41 tok/s. For image generation, SDXL runs at 17.29 it/s, and FLUX.1-dev at 4.23 it/s. All 12 workloads fit in 80GB. There is no model in our suite this card has to turn down. NVIDIA H100 80GB HBM3 isn't a retail purchase for most people. It's rented by the hour. You can run this exact card on RunPod.
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| SmolLM2-135M | 898.53 tok/s | 125 W39°CQ4_K_M | ✓ Measured |
| gemma-3-270m | 961.29 tok/s | 128 W39°CQ4_K_M | ✓ Measured |
| SmolLM2-360M | 787.64 tok/s | 138 W43°CQ4_K_M | ✓ Measured |
| Qwen1.5-0.5B | 847.02 tok/s | 134 W39°CQ4_K_M | ✓ Measured |
| Qwen2 0.5B | 881.94 tok/s | 147 W44°CQ4_K_M | ✓ Measured |
| Qwen2.5-0.5B | 890 tok/s | 134 W38°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder-0.5B | 892.81 tok/s | 130 W39°CQ4_K_M | ✓ Measured |
| Qwen3 0.6B | 713.53 tok/s | 104 W46°CQ4_K_M | ✓ Measured |
| Qwen3-0.6B | 710.97 tok/s | 140 W40°CQ4_K_M | ✓ Measured |
| Llama 3.2 1B | 880.64 tok/s | 188 W48°CQ4_K_M | ✓ Measured |
| gemma-3-1b | 504.17 tok/s | 152 W40°CQ4_K_M | ✓ Measured |
| LFM2.5-1.2B | 926.61 tok/s | 138 W47°CQ4_K_M | ✓ Measured |
| DeepSeek-R1 Distill 1.5B | 537.32 tok/s | 213 W47°CQ4_K_M | ✓ Measured |
| Qwen2-1.5B | 536.52 tok/s | 162 W41°CQ4_K_M | ✓ Measured |
| Qwen2.5-1.5B | 537.24 tok/s | 166 W41°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder-1.5B | 538.44 tok/s | 165 W46°CQ4_K_M | ✓ Measured |
| Qwen3 1.7B | 556.48 tok/s | 177 W49°CQ4_K_M | ✓ Measured |
| Qwen3-1.7B | 557.09 tok/s | 162 W41°CQ4_K_M | ✓ Measured |
| MiniCPM5 2B | 428.69 tok/s | 164 W53°CQ4_K_M | ✓ Measured |
| gemma-2-2b | 392.26 tok/s | 162 W43°CQ4_K_M | ✓ Measured |
| gemma-2-2b-it-abliterated | 392.65 tok/s | 167 W44°CQ4_K_M | ✓ Measured |
| LFM2.5 2.6B | 482.59 tok/s | 160 W53°CQ4_K_M | ✓ Measured |
| AI21-Jamba-Reasoning-3B | 363.09 tok/s | 179 W42°CQ4_K_M | ✓ Measured |
| Granite 4.1 3B | 332.28 tok/s | 196 W44°CQ4_K_M | ✓ Measured |
| Hermes-3-Llama-3.2-3B | 423.72 tok/s | 167 W43°CQ4_K_M | ✓ Measured |
| Llama 3.2 3B | 422.87 tok/s | 194 W50°CQ4_K_M | ✓ Measured |
| Llama-3.2-3B-Instruct-uncensored | 424.25 tok/s | 193 W44°CQ4_K_M | ✓ Measured |
| Nanbeige4.2-3B | ✕ Won't fit needs ~4 GB | VRAM-gated at this precision | ✓ Measured |
| Qwen2.5-3B | 395.52 tok/s | 168 W43°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder-3B | 394 tok/s | 201 W48°CQ4_K_M | ✓ Measured |
| SmolLM3 3B | 397.99 tok/s | 230 W49°CQ4_K_M | ✓ Measured |
| SmolLM3-3B | 399.14 tok/s | 195 W41°CQ4_K_M | ✓ Measured |
| Phi-4 Mini 3.8B | 388.86 tok/s | 210 W51°CQ4_K_M | ✓ Measured |
| Gemma 3 4B | 289.61 tok/s | 211 W49°CQ4_K_M | ✓ Measured |
| Nemotron 3 Nano 4B | 381.9 tok/s | 180 W53°CQ4_K_M | ✓ Measured |
| Qwen3 4B | 310.26 tok/s | 3.3 GB peak169 W45°C1.84 tok/WQ4_K_M | ✓ Measured |
| Qwen3-4B-Instruct-2507 | 311.47 tok/s | 191 W44°CQ4_K_M | ✓ Measured |
| Qwen3-4B-Instruct-2507 | 310.58 tok/s | 189 W44°CQ4_K_M | ✓ Measured |
| Qwen3-4B-Thinking-2507 | 310.84 tok/s | 201 W43°CQ4_K_M | ✓ Measured |
| phi-2 | 340.09 tok/s | 189 W42°CQ4_K_M | ✓ Measured |
| Dolphin X1 Trinity Nano 6B | 247.94 tok/s | 141 W38°CQ4_K_M | ✓ Measured |
| Phi-3.5-mini | 345.07 tok/s | 193 W43°CQ4_K_M | ✓ Measured |
| Phi-4-mini | 381.52 tok/s | 188 W49°CQ4_K_M | ✓ Measured |
| DeepSeek Coder 7B Instruct v1.5 | 293.91 tok/s | 224 W48°CQ4_K_M | ✓ Measured |
| DeepSeek-R1 Distill 7B | 264.42 tok/s | 213 W37°CQ4_K_M | ✓ Measured |
| Llama-2-7B | 284.62 tok/s | 225 W49°CQ4_K_M | ✓ Measured |
| Mistral 7B v0.3 | 275.85 tok/s | 235 W53°CQ4_K_M | ✓ Measured |
| Mistral-7B-Instruct-v0.1 | 276.29 tok/s | 260 W49°CQ4_K_M | ✓ Measured |
| Mistral-7B-Instruct-v0.2 | 276.73 tok/s | 265 W50°CQ4_K_M | ✓ Measured |
| Mistral-7B-Instruct-v0.3 | 276.61 tok/s | 239 W43°CQ4_K_M | ✓ Measured |
| Qwen2.5-7B | 264.17 tok/s | 240 W46°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder 7B | 263.91 tok/s | 241 W52°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder-7B-Instruct-abliterated | 264.24 tok/s | 239 W46°CQ4_K_M | ✓ Measured |
| Qwen2.5-VL 7B Instruct | 267.72 tok/s | 227 W44°CQ4_K_M | ✓ Measured |
| DarkIdol-Llama-3.1-8B-Instruct-1.2-Uncensored | 262.76 tok/s | 232 W49°CQ4_K_M | ✓ Measured |
| DeepSeek-R1 Distill Llama 8B | 261.15 tok/s | 228 W37°CQ4_K_M | ✓ Measured |
| DeepSeek-R1-0528-Qwen3-8B | 245.26 tok/s | 260 W46°CQ4_K_M | ✓ Measured |
| Dolphin 3.0 Llama 3.1 8B | 266.19 tok/s | 107 W41°CQ4_K_M | ✓ Measured |
| Dolphin X1 8B | 266 tok/s | 183 W41°CQ4_K_M | ✓ Measured |
| Josiefied-Qwen3-8B-abliterated-v1 | 245 tok/s | 228 W46°CQ4_K_M | ✓ Measured |
| L3-8B-Stheno-v3.2 | 261.83 tok/s | 242 W47°CQ4_K_M | ✓ Measured |
| LFM2.5-8B-A1B | 581.98 tok/s | 156 W41°CQ4_K_M | ✓ Measured |
| Llama 3 8B | 264.36 tok/s | 258 W48°CQ4_K_M | ✓ Measured |
| Llama 3.1 8B | 261.83 tok/s | 5.2 GB peak202 W49°C1.3 tok/WQ4_K_M | ✓ Measured |
| Meta-Llama-3.1-8B | 261.7 tok/s | 228 W46°CQ4_K_M | ✓ Measured |
| Qwen3 8B | 244.22 tok/s | 174 W36°CQ4_K_M | ✓ Measured |
| Qwen3-8B | 244.2 tok/s | 243 W51°CQ4_K_M | ✓ Measured |
| dolphin-2.9-llama3-8b | 262.47 tok/s | 233 W47°CQ4_K_M | ✓ Measured |
| Nemotron Nano 9B v2 | 227.65 tok/s | 269 W49°CQ4_K_M | ✓ Measured |
| Ornith 1.5 9B | 221.43 tok/s | 246 W55°CQ4_K_M | ✓ Measured |
| Ornith-1.0-9B | 216.89 tok/s | 239 W46°CQ4_K_M | ✓ Measured |
| gemma-2-9b | 174.7 tok/s | 255 W48°CQ4_K_M | ✓ Measured |
| Gemma 3 12B | 150.58 tok/s | 234 W38°CQ4_K_M | ✓ Measured |
| Gemma 3 12B (Q3_K_M) | 125.92 tok/s | 282 W46°CQ3_K_M | ✓ Measured |
| Gemma 4 12B | 148.17 tok/s | 244 W37°CQ4_K_M | ✓ Measured |
| NemoMix-Unleashed-12B | 177.28 tok/s | 259 W47°CQ4_K_M | ✓ Measured |
| DeepSeek-R1 Distill 14B | 144.82 tok/s | 252 W40°CQ4_K_M | ✓ Measured |
| DeepSeek-R1 Distill 14B (Q3_K_M) | 118.43 tok/s | 296 W48°CQ3_K_M | ✓ Measured |
| EVA-Qwen2.5-14B-v0.2 | 145.15 tok/s | 307 W50°CQ4_K_M | ✓ Measured |
| Hermes-4-14B | 152 tok/s | 285 W48°CQ4_K_M | ✓ Measured |
| Phi-4 14B | 165.17 tok/s | 238 W40°CQ4_K_M | ✓ Measured |
| Phi-4 14B (Q3_K_M) | 138.87 tok/s | 310 W48°CQ3_K_M | ✓ Measured |
| Qwen2.5-14B | 144.7 tok/s | 286 W48°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder 14B | 144.84 tok/s | 9 GB peak233 W51°C0.62 tok/WQ4_K_M | ✓ Measured |
| Qwen2.5-Coder-14B-Instruct-abliterated | 145.01 tok/s | 296 W48°CQ4_K_M | ✓ Measured |
| Qwen3 14B | 151.16 tok/s | 208 W39°CQ4_K_M | ✓ Measured |
| Qwen3-14B | 152.2 tok/s | 280 W46°CQ4_K_M | ✓ Measured |
| Uncensored | 144.92 tok/s | 279 W48°CQ4_K_M | ✓ Measured |
| StarCoder2 15B | 134.92 tok/s | 181 W43°CQ4_K_M | ✓ Measured |
| Mistral-Nemo-Instruct-2407 | 177.38 tok/s | 280 W47°CQ4_K_M | ✓ Measured |
| gpt-oss-20b | 346.8 tok/s | 165 W43°CQ4_K_M | ✓ Measured |
| Codestral 22B | 110.05 tok/s | 224 W44°CQ4_K_M | ✓ Measured |
| Codestral 22B (Q3_K_M) | 87.21 tok/s | 338 W50°CQ3_K_M | ✓ Measured |
| GLM-4.7-Flash-REAP-23B-A3B | 168.41 tok/s | 172 W41°CQ4_K_M | ✓ Measured |
| DeepSeek-Coder-V2-Lite | 309.14 tok/s | 173 W44°CQ4_K_M | ✓ Measured |
| Cydonia-24B-v4.3 | 105.49 tok/s | 311 W50°CQ4_K_M | ✓ Measured |
| Devstral Small 24B | 107.79 tok/s | 140 W44°CQ4_K_M | ✓ Measured |
| Dolphin 3.0 R1 Mistral 24B | 107.83 tok/s | 124 W45°CQ4_K_M | ✓ Measured |
| Dolphin Mistral 24B Venice | 109.13 tok/s | 191 W48°CQ4_K_M | ✓ Measured |
| Dolphin-Mistral-24B-Venice-Edition | 105.63 tok/s | 323 W50°CQ4_K_M | ✓ Measured |
| Mistral Small 24B | 105.62 tok/s | 241 W41°CQ4_K_M | ✓ Measured |
| Mistral Small 24B (Q3_K_M) | 84 tok/s | 341 W50°CQ3_K_M | ✓ Measured |
| Gemma 4 26B A4B | 218.39 tok/s | 169 W52°CQ4_K_M | ✓ Measured |
| Gemma 3 27B | 84.77 tok/s | 216 W50°CQ4_K_M | ✓ Measured |
| Qwen3.6 27B | 79.74 tok/s | 301 W59°CQ4_K_M | ✓ Measured |
| Qwen3.8 27B | 80.5 tok/s | 303 W59°CQ4_K_M | ✓ Measured |
| Nemotron 3.5 Lightning 30B A3B | 323.62 tok/s | 162 W45°CQ4_K_M | ✓ Measured |
| Nemotron-3-Nano-30B-A3B | 316.51 tok/s | 158 W43°CQ4_K_M | ✓ Measured |
| Qwen3 30B A3B | 283.78 tok/s | 142 W34°CQ4_K_M | ✓ Measured |
| Qwen3 30B A3B (Q3_K_M) | 242.6 tok/s | 175 W41°CQ3_K_M | ✓ Measured |
| Qwen3 30B A3B Instruct 2507 | 299.76 tok/s | 166 W43°CQ4_K_M | ✓ Measured |
| Qwen3-30B-A3B | 285.44 tok/s | 160 W42°CQ4_K_M | ✓ Measured |
| Qwen3-Coder 30B A3B | 296.22 tok/s | 127 W37°CQ4_K_M | ✓ Measured |
| Gemma 4 31B | 76.59 tok/s | 324 W59°CQ4_K_M | ✓ Measured |
| DeepSeek-R1 Distill 32B | 75.8 tok/s | 210 W52°CQ4_K_M | ✓ Measured |
| DeepSeek-R1-Distill-Qwen-32B-abliterated | 73.48 tok/s | 344 W50°CQ4_K_M | ✓ Measured |
| Olmo-3.1-32B-Think | 75.53 tok/s | 344 W52°CQ4_K_M | ✓ Measured |
| QwQ 32B | 75.82 tok/s | 213 W51°CQ4_K_M | ✓ Measured |
| Qwen2.5-32B | 73.5 tok/s | 339 W51°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder 32B | 75.81 tok/s | 175 W52°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder 32B (Q3_K_M) | 58.05 tok/s | 351 W50°CQ3_K_M | ✓ Measured |
| Qwen3 32B | 74.07 tok/s | 19 GB peak218 W54°C0.34 tok/WQ4_K_M | ✓ Measured |
| Dolphin 2.9.1 Yi 1.5 34B | 76.76 tok/s | 199 W51°CQ4_K_M | ✓ Measured |
| Ornith 1.5 35B A3B | 240.25 tok/s | 165 W41°CQ4_K_M | ✓ Measured |
| Ornith-1.0-35B | 221.65 tok/s | 163 W42°CQ4_K_M | ✓ Measured |
| Qwen-AgentWorld-35B-A3B | 219.1 tok/s | 168 W43°CQ4_K_M | ✓ Measured |
| Qwen3.6 35B A3B | 236.68 tok/s | 163 W52°CQ4_K_M | ✓ Measured |
| GLM-4.7-Flash | 186.05 tok/s | 169 W42°CQ4_K_M | ✓ Measured |
| Laguna-XS-2.1 | ✕ Won't fit needs ~26 GB | VRAM-gated at this precision | ✓ Measured |
| KAT-Coder-V2.5-Dev | 231.83 tok/s | 166 W42°CQ4_K_M | ✓ Measured |
| DeepSeek-R1-Distill-Llama-70B | 41.17 tok/s | 380 W56°CQ4_K_M | ✓ Measured |
| Hermes-4-70B | 41.2 tok/s | 370 W55°CQ4_K_M | ✓ Measured |
| Llama 3.3 70B | 41 tok/s | 40.1 GB peak293 W57°C0.14 tok/WQ4_K_M | ✓ Measured |
| Llama-3.3-70B-Instruct-abliterated | 41.15 tok/s | 357 W56°CQ4_K_M | ✓ Measured |
| Meta-Llama-3.1-70B | 41.19 tok/s | 379 W52°CQ4_K_M | ✓ Measured |
| Qwen2.5-72B | 40.78 tok/s | 366 W56°CQ4_K_M | ✓ Measured |
| Qwen3-Next-80B-A3B | 186.22 tok/s | 148 W40°CQ4_K_M | ✓ Measured |
| Qwen3-Next-80B-A3B-Thinking | 191.08 tok/s | 148 W40°CQ4_K_M | ✓ Measured |
| Qwen3-Next-80B-A3B-Thinking | 188.14 tok/s | 151 W42°CQ4_K_M | ✓ Measured |
| Qwen3-Coder-Next | 193.79 tok/s | 149 W42°CQ4_K_M | ✓ Measured |
| Qwen3-Coder-Next | 188.21 tok/s | 152 W41°CQ4_K_M | ✓ Measured |
| Qwen3-Coder-Next-abliterated | 190.32 tok/s | 152 W41°CQ4_K_M | ✓ Measured |
| gpt-oss-120b | 220.43 tok/s | 154 W55°CMXFP4 | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Stable Diffusion 1.5 | 80.45 images/min | 294 W31°C | ✓ Measured |
| SDXL Turbo | 618.27 images/min | 120 W25°C | ✓ Measured |
| Stable Diffusion XL | 34.58 images/min | 16.2 GB peak539 W56°C1.7 s/img | ✓ Measured |
| Playground v2.5 | 21.48 images/min | 543 W42°C | ✓ Measured |
| FLUX.2 klein 4B | 108.33 images/min | 512 W53°C | ✓ Measured |
| Z-Image | 3.58 images/min | 689 W53°C | ✓ Measured |
| Z-Image Turbo (1024px) | 39.42 images/min | 640 W45°C | ✓ Measured |
| Z-Image Turbo | 22.13 images/min | 26 GB peak661 W59°C2.7 s/img | ✓ Measured |
| FLUX.1 Schnell | 58.36 images/min | 614 W45°C | ✓ Measured |
| FLUX.1 dev | 9.06 images/min | 36.9 GB peak662 W63°C6.6 s/img | ✓ Measured |
| FLUX.2 klein 9B | 64.87 images/min | 649 W59°C | ✓ Measured |
| Qwen-Image 2.1 | 4.59 images/min | 690 W68°C | ✓ Measured |
| Krea 2 Turbo | 6.85 images/min | 638 W60°C | ✓ Measured |
| Qwen-Image | 3 images/min | 679 W63°C | ✓ Measured |
| Qwen-Image 2512 | 3 images/min | 669 W60°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| FLUX.1 Kontext dev | 4.26 images/min | 35.7 GB peak683 W65°C14.1 s/img | ✓ Measured |
| Qwen-Image-Edit 2509 | 1.65 images/min | 689 W65°C | ✓ Measured |
| Qwen-Image-Edit | 3.68 images/min | 60.4 GB peak683 W65°C16.3 s/img | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| BiRefNet | 1310.97 images/min | 123 W33°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Swin2SR 4x Upscaler | 25.48 images/min | 203 W38°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Stable Video Diffusion | 4.88 clips/min | 615 W64°C | ✓ Measured |
| LTX-Video (image to video) | 10.62 clips/min | 595 W60°C | ✓ Measured |
| Wan 2.2 TI2V-5B (image to video) | 4.13 clips/min | 673 W64°C | ✓ Measured |
| Stable Video Diffusion XT | 2.88 clips/min | 619 W65°C | ✓ Measured |
| CogVideoX-5B I2V | 0.88 clips/min | 657 W67°C | ✓ Measured |
| Wan 2.1 I2V 14B (480p) | 0.4 clips/min | 692 W67°C | ✓ Measured |
| Wan 2.2 I2V A14B | 0.4 clips/min | 692 W81°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| CogVideoX-2B | 1.17 frames/s | 666 W74°C42 s/clip | ✓ Measured |
| Mochi 1 Preview | 0.38 frames/s | 683 W69°C128.1 s/clip | ✓ Measured |
| LTX-Video (distilled) | 17.47 frames/s | 60.3 GB peak554 W60°C5.6 s/clip | ✓ Measured |
| Wan 2.2 5B (720p) | 1.33 frames/s | 61 GB peak669 W67°C36.8 s/clip | ✓ Measured |
| Wan 2.2 T2V A14B | 0.34 frames/s | 689 W77°C146.1 s/clip | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| TripoSR Image-to-3D | 2107.4 assets/hour | 171 W39°C | ✓ Measured |
| TripoSG Image-to-3D | 559.4 assets/hour | 529 W70°C | ✓ Measured |
| TRELLIS Image-to-3D | 167.2 assets/hour | ✓ Measured | |
| TRELLIS.2 Image-to-3D | 87.9 assets/hour | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| MusicGen Small | 2.47 x realtime | 163 W37°C | ✓ Measured |
| ACE-Step 1.5 | 27.05 x realtime | 179 W32°C | ✓ Measured |
| ACE-Step v1 3.5B | 14.83 x realtime | 237 W34°C | ✓ Measured |
| DiffRhythm 2 | 5.63 x realtime | 170 W38°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| EzAudio XL | 1.58 x realtime | 474 W41°C | ✓ Measured |
| MOSS-SoundEffect v2.0 | 2.8 x realtime | 577 W57°C | ✓ Measured |
| MiDashengLM-Gen | 1.43 x realtime | 230 W38°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Whisper large-v3 | 181.47 x realtime | 138 W37°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Kokoro TTS 82M | 222.02 x realtime | 126 W37°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Depth Anything V2 Small | 1081.54 images/min | 116 W33°C | ✓ Measured |
| Depth Anything V2 Large | 918.71 images/min | 117 W34°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| SAM ViT-Base | 1431.27 images/min | 130 W38°C | ✓ Measured |
| SAM ViT-Huge | 320.64 images/min | 285 W47°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Florence-2 Base | 249.62 images/min | 125 W35°C | ✓ Measured |
| Florence-2 Large | 156.15 images/min | 143 W36°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| TinyLlama 1.1B LoRA | 17522 train tok/s | 301 W46°C | ✓ Measured |
| Qwen2.5 1.5B LoRA | 16034.9 train tok/s | 335 W47°C | ✓ Measured |
| SmolLM2 1.7B LoRA | 18910.3 train tok/s | 398 W50°C | ✓ Measured |
| Qwen2.5 7B LoRA | 8514.3 train tok/s | 633 W56°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| TinyLlama 1.1B served | 8336.7 serve tok/s | 197 W35°C | ✓ Measured |
| Qwen2.5 1.5B served | 6794.8 serve tok/s | 237 W37°C | ✓ Measured |
| SmolLM2 1.7B served | 6509.9 serve tok/s | 262 W39°C | ✓ Measured |
| Qwen2.5 7B served | 3601.1 serve tok/s | 424 W42°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Qwen2.5 1.5B + 0.5B draft | 0.73 x vs solo | ✓ Measured |
| Architecture | Hopper |
| CUDA cores | 16,896 |
| VRAM | 80GB HBM3 |
| Memory bus | 5120-bit |
| Memory bandwidth | 3350 GB/s |
| Boost clock | 1,980 MHz |
| TDP | 700 W |
| Process | TSMC 4N |
| Interface | SXM5 |
| Release date | 2022-09-20 |
| Launch MSRP | $30,000 |
NVIDIA H100 80GB HBM3 scores 62.6/100, #6 of 102. It ran all 12 workloads. Every figure here is our own measurement.
100% = this card, AI & Machine Learning headline metric (AI Score). #6 of 21 datacenter cards in this vertical.
| GPU | Relative | % | AI Score |
|---|---|---|---|
| NVIDIA B200 | 125% | 78 | |
| NVIDIA B100 | 107% | 67 | |
| NVIDIA GH200 Grace Hopper | 105% | 65.6 | |
| NVIDIA H200 | 104% | 65 | |
| NVIDIA H100 80GB HBM3 | 100% | 62.6 | |
| NVIDIA H800 80GB | 100% | 62.6 | |
| NVIDIA H100 NVL | 93% | 58.1 | |
| NVIDIA RTX PRO 6000 Blackwell Server Edition | 80% | 49.9 | |
| NVIDIA H100 PCIe | 80% | 49.8 |
← All AI & Machine Learning GPU rankings
| Transistors | 80,000 million |
| Die size | 814 mm² |
| Process node | 4 nm |
| Fabricated by | TSMC |
| Transistor density | 98.3 million per mm² |
Denser than 90% of the 746 cards we have silicon data for. Density is the clearest measure of what a process node bought: a card that gained it without growing the die got its speed from the fab rather than the architecture.
Silicon figures from Wikipedia (CC BY-SA 4.0). Benchmarks on this page are our own. Compare every chip.
Whole-job timings, composed from our measured per-model results on this card.
| Workflow | Time | Energy | Basis |
|---|---|---|---|
| Depth pass on a batch | 5 s | 0.11 Wh | measured |
| Voiceovers from scripts | 17 s | 0.28 Wh | measured |
| Podcast episode pass | 34 s | 1.05 Wh | all 3 stages measured |
| Masking run | 1.6 min | 7.4 Wh | measured |
| Transcribe and subtitle videos | 1.7 min | 3.79 Wh | measured |
| 24-frame storyboard | 1.8 min | 13.08 Wh | all 2 stages measured |
| 60-second AI short film | 2.5 min | 17.46 Wh | all 3 stages measured |
| 6-panel comic page | 3 min | 23.89 Wh | all 3 stages measured |
| Character sheet, 12 poses | 3.7 min | 33.26 Wh | all 2 stages measured |
| Animate a batch of images | 5.5 min | 54.4 Wh | measured |
| Caption a training dataset | 6.5 min | 15.22 Wh | measured |
| Full codebase review | 6.9 min | 26.8 Wh | measured |
| Short social clips | 8 min | 73.89 Wh | all 3 stages measured |
| Product catalogue cutout | 8.1 min | 26.87 Wh | all 2 stages measured |
| 3D game asset kit | 8.6 min | 5.19 Wh | all 2 stages measured |
| Product photo shoot | 11.1 min | 117.19 Wh | all 2 stages measured |
| Long-form article batch | 11.4 min | 55.51 Wh | measured |
| Product shoot, start to finish | 12.9 min | 122.56 Wh | all 4 stages measured |
| Photos to 3D models | 18.9 min | 0.04 Wh | all 2 stages measured |
| Photo restoration batch | 23.8 min | 267 Wh | measured |
| Restore and enlarge photos | 27.8 min | 280.28 Wh | all 2 stages measured |
This card is $30,000 to buy. The cheapest listed rate on Vast.ai is $2.003/hour, but that is the floor: we budget $2.404/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 12,481 GPU-hours. Below it you are paying for idle silicon.
| How you would use it | GPU-hours a year | Rental cost a year | Time to break even |
|---|---|---|---|
| 2 hours a day, hobby | 730 | $1,755 | 17.1 years |
| 8 hours a day, working on it | 2,920 | $7,019 | 4.3 years |
| 24/7, always-on agent | 8,760 | $21,056 | 1.4 years |
At hobby usage this card is very unlikely to pay for itself before it is superseded. Rent it. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.
Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.