NVIDIA H100 80GB HBM3, AI & Machine Learning Benchmarks & Specs

80GB · AI Score 62.6/100 · first-party measured on 12 AI workloads

62.6 AI Score ✓ Measured

Every number on this page is first-party: NVIDIA H100 80GB HBM3 was run on our pinned 12-workload AI suite on 2026-07-10, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) NVIDIA H100 80GB HBM3 delivers about 261.83 tokens/sec. Stepping up to Qwen3 32B it holds roughly 74.07 tok/s. The full Llama 3.3 70B still runs, at about 41 tok/s. For image generation, SDXL runs at 17.29 it/s, and FLUX.1-dev at 4.23 it/s. All 12 workloads fit in 80GB. There is no model in our suite this card has to turn down. NVIDIA H100 80GB HBM3 isn't a retail purchase for most people. It's rented by the hour. You can run this exact card on RunPod.

AI & Machine Learning benchmark results

Text Generation tok/s 143

gemma-3-270m961.29
LFM2.5-1.2B926.61
SmolLM2-135M898.53
Qwen2.5-Coder-0.5B892.81
Qwen2.5-0.5B890
Qwen2 0.5B881.94
Llama 3.2 1B880.64
Qwen1.5-0.5B847.02
SmolLM2-360M787.64
Qwen3 0.6B713.53
Qwen3-0.6B710.97
LFM2.5-8B-A1B581.98
WorkloadResultTelemetryData
SmolLM2-135M898.53 tok/s
125 W39°CQ4_K_M
✓ Measured
gemma-3-270m961.29 tok/s
128 W39°CQ4_K_M
✓ Measured
SmolLM2-360M787.64 tok/s
138 W43°CQ4_K_M
✓ Measured
Qwen1.5-0.5B847.02 tok/s
134 W39°CQ4_K_M
✓ Measured
Qwen2 0.5B881.94 tok/s
147 W44°CQ4_K_M
✓ Measured
Qwen2.5-0.5B890 tok/s
134 W38°CQ4_K_M
✓ Measured
Qwen2.5-Coder-0.5B892.81 tok/s
130 W39°CQ4_K_M
✓ Measured
Qwen3 0.6B713.53 tok/s
104 W46°CQ4_K_M
✓ Measured
Qwen3-0.6B710.97 tok/s
140 W40°CQ4_K_M
✓ Measured
Llama 3.2 1B880.64 tok/s
188 W48°CQ4_K_M
✓ Measured
gemma-3-1b504.17 tok/s
152 W40°CQ4_K_M
✓ Measured
LFM2.5-1.2B926.61 tok/s
138 W47°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 1.5B537.32 tok/s
213 W47°CQ4_K_M
✓ Measured
Qwen2-1.5B536.52 tok/s
162 W41°CQ4_K_M
✓ Measured
Qwen2.5-1.5B537.24 tok/s
166 W41°CQ4_K_M
✓ Measured
Qwen2.5-Coder-1.5B538.44 tok/s
165 W46°CQ4_K_M
✓ Measured
Qwen3 1.7B556.48 tok/s
177 W49°CQ4_K_M
✓ Measured
Qwen3-1.7B557.09 tok/s
162 W41°CQ4_K_M
✓ Measured
MiniCPM5 2B428.69 tok/s
164 W53°CQ4_K_M
✓ Measured
gemma-2-2b392.26 tok/s
162 W43°CQ4_K_M
✓ Measured
gemma-2-2b-it-abliterated392.65 tok/s
167 W44°CQ4_K_M
✓ Measured
LFM2.5 2.6B482.59 tok/s
160 W53°CQ4_K_M
✓ Measured
AI21-Jamba-Reasoning-3B363.09 tok/s
179 W42°CQ4_K_M
✓ Measured
Granite 4.1 3B332.28 tok/s
196 W44°CQ4_K_M
✓ Measured
Hermes-3-Llama-3.2-3B423.72 tok/s
167 W43°CQ4_K_M
✓ Measured
Llama 3.2 3B422.87 tok/s
194 W50°CQ4_K_M
✓ Measured
Llama-3.2-3B-Instruct-uncensored424.25 tok/s
193 W44°CQ4_K_M
✓ Measured
Nanbeige4.2-3B✕ Won't fit needs ~4 GBVRAM-gated at this precision✓ Measured
Qwen2.5-3B395.52 tok/s
168 W43°CQ4_K_M
✓ Measured
Qwen2.5-Coder-3B394 tok/s
201 W48°CQ4_K_M
✓ Measured
SmolLM3 3B397.99 tok/s
230 W49°CQ4_K_M
✓ Measured
SmolLM3-3B399.14 tok/s
195 W41°CQ4_K_M
✓ Measured
Phi-4 Mini 3.8B388.86 tok/s
210 W51°CQ4_K_M
✓ Measured
Gemma 3 4B289.61 tok/s
211 W49°CQ4_K_M
✓ Measured
Nemotron 3 Nano 4B381.9 tok/s
180 W53°CQ4_K_M
✓ Measured
Qwen3 4B310.26 tok/s
3.3 GB peak169 W45°C1.84 tok/WQ4_K_M
✓ Measured
Qwen3-4B-Instruct-2507311.47 tok/s
191 W44°CQ4_K_M
✓ Measured
Qwen3-4B-Instruct-2507310.58 tok/s
189 W44°CQ4_K_M
✓ Measured
Qwen3-4B-Thinking-2507310.84 tok/s
201 W43°CQ4_K_M
✓ Measured
phi-2340.09 tok/s
189 W42°CQ4_K_M
✓ Measured
Dolphin X1 Trinity Nano 6B247.94 tok/s
141 W38°CQ4_K_M
✓ Measured
Phi-3.5-mini345.07 tok/s
193 W43°CQ4_K_M
✓ Measured
Phi-4-mini381.52 tok/s
188 W49°CQ4_K_M
✓ Measured
DeepSeek Coder 7B Instruct v1.5293.91 tok/s
224 W48°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 7B264.42 tok/s
213 W37°CQ4_K_M
✓ Measured
Llama-2-7B284.62 tok/s
225 W49°CQ4_K_M
✓ Measured
Mistral 7B v0.3275.85 tok/s
235 W53°CQ4_K_M
✓ Measured
Mistral-7B-Instruct-v0.1276.29 tok/s
260 W49°CQ4_K_M
✓ Measured
Mistral-7B-Instruct-v0.2276.73 tok/s
265 W50°CQ4_K_M
✓ Measured
Mistral-7B-Instruct-v0.3276.61 tok/s
239 W43°CQ4_K_M
✓ Measured
Qwen2.5-7B264.17 tok/s
240 W46°CQ4_K_M
✓ Measured
Qwen2.5-Coder 7B263.91 tok/s
241 W52°CQ4_K_M
✓ Measured
Qwen2.5-Coder-7B-Instruct-abliterated264.24 tok/s
239 W46°CQ4_K_M
✓ Measured
Qwen2.5-VL 7B Instruct267.72 tok/s
227 W44°CQ4_K_M
✓ Measured
DarkIdol-Llama-3.1-8B-Instruct-1.2-Uncensored262.76 tok/s
232 W49°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill Llama 8B261.15 tok/s
228 W37°CQ4_K_M
✓ Measured
DeepSeek-R1-0528-Qwen3-8B245.26 tok/s
260 W46°CQ4_K_M
✓ Measured
Dolphin 3.0 Llama 3.1 8B266.19 tok/s
107 W41°CQ4_K_M
✓ Measured
Dolphin X1 8B266 tok/s
183 W41°CQ4_K_M
✓ Measured
Josiefied-Qwen3-8B-abliterated-v1245 tok/s
228 W46°CQ4_K_M
✓ Measured
L3-8B-Stheno-v3.2261.83 tok/s
242 W47°CQ4_K_M
✓ Measured
LFM2.5-8B-A1B581.98 tok/s
156 W41°CQ4_K_M
✓ Measured
Llama 3 8B264.36 tok/s
258 W48°CQ4_K_M
✓ Measured
Llama 3.1 8B261.83 tok/s
5.2 GB peak202 W49°C1.3 tok/WQ4_K_M
✓ Measured
Meta-Llama-3.1-8B261.7 tok/s
228 W46°CQ4_K_M
✓ Measured
Qwen3 8B244.22 tok/s
174 W36°CQ4_K_M
✓ Measured
Qwen3-8B244.2 tok/s
243 W51°CQ4_K_M
✓ Measured
dolphin-2.9-llama3-8b262.47 tok/s
233 W47°CQ4_K_M
✓ Measured
Nemotron Nano 9B v2227.65 tok/s
269 W49°CQ4_K_M
✓ Measured
Ornith 1.5 9B221.43 tok/s
246 W55°CQ4_K_M
✓ Measured
Ornith-1.0-9B216.89 tok/s
239 W46°CQ4_K_M
✓ Measured
gemma-2-9b174.7 tok/s
255 W48°CQ4_K_M
✓ Measured
Gemma 3 12B150.58 tok/s
234 W38°CQ4_K_M
✓ Measured
Gemma 3 12B (Q3_K_M)125.92 tok/s
282 W46°CQ3_K_M
✓ Measured
Gemma 4 12B148.17 tok/s
244 W37°CQ4_K_M
✓ Measured
NemoMix-Unleashed-12B177.28 tok/s
259 W47°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 14B144.82 tok/s
252 W40°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 14B (Q3_K_M)118.43 tok/s
296 W48°CQ3_K_M
✓ Measured
EVA-Qwen2.5-14B-v0.2145.15 tok/s
307 W50°CQ4_K_M
✓ Measured
Hermes-4-14B152 tok/s
285 W48°CQ4_K_M
✓ Measured
Phi-4 14B165.17 tok/s
238 W40°CQ4_K_M
✓ Measured
Phi-4 14B (Q3_K_M)138.87 tok/s
310 W48°CQ3_K_M
✓ Measured
Qwen2.5-14B144.7 tok/s
286 W48°CQ4_K_M
✓ Measured
Qwen2.5-Coder 14B144.84 tok/s
9 GB peak233 W51°C0.62 tok/WQ4_K_M
✓ Measured
Qwen2.5-Coder-14B-Instruct-abliterated145.01 tok/s
296 W48°CQ4_K_M
✓ Measured
Qwen3 14B151.16 tok/s
208 W39°CQ4_K_M
✓ Measured
Qwen3-14B152.2 tok/s
280 W46°CQ4_K_M
✓ Measured
Uncensored144.92 tok/s
279 W48°CQ4_K_M
✓ Measured
StarCoder2 15B134.92 tok/s
181 W43°CQ4_K_M
✓ Measured
Mistral-Nemo-Instruct-2407177.38 tok/s
280 W47°CQ4_K_M
✓ Measured
gpt-oss-20b346.8 tok/s
165 W43°CQ4_K_M
✓ Measured
Codestral 22B110.05 tok/s
224 W44°CQ4_K_M
✓ Measured
Codestral 22B (Q3_K_M)87.21 tok/s
338 W50°CQ3_K_M
✓ Measured
GLM-4.7-Flash-REAP-23B-A3B168.41 tok/s
172 W41°CQ4_K_M
✓ Measured
DeepSeek-Coder-V2-Lite309.14 tok/s
173 W44°CQ4_K_M
✓ Measured
Cydonia-24B-v4.3105.49 tok/s
311 W50°CQ4_K_M
✓ Measured
Devstral Small 24B107.79 tok/s
140 W44°CQ4_K_M
✓ Measured
Dolphin 3.0 R1 Mistral 24B107.83 tok/s
124 W45°CQ4_K_M
✓ Measured
Dolphin Mistral 24B Venice109.13 tok/s
191 W48°CQ4_K_M
✓ Measured
Dolphin-Mistral-24B-Venice-Edition105.63 tok/s
323 W50°CQ4_K_M
✓ Measured
Mistral Small 24B105.62 tok/s
241 W41°CQ4_K_M
✓ Measured
Mistral Small 24B (Q3_K_M)84 tok/s
341 W50°CQ3_K_M
✓ Measured
Gemma 4 26B A4B218.39 tok/s
169 W52°CQ4_K_M
✓ Measured
Gemma 3 27B84.77 tok/s
216 W50°CQ4_K_M
✓ Measured
Qwen3.6 27B79.74 tok/s
301 W59°CQ4_K_M
✓ Measured
Qwen3.8 27B80.5 tok/s
303 W59°CQ4_K_M
✓ Measured
Nemotron 3.5 Lightning 30B A3B323.62 tok/s
162 W45°CQ4_K_M
✓ Measured
Nemotron-3-Nano-30B-A3B316.51 tok/s
158 W43°CQ4_K_M
✓ Measured
Qwen3 30B A3B283.78 tok/s
142 W34°CQ4_K_M
✓ Measured
Qwen3 30B A3B (Q3_K_M)242.6 tok/s
175 W41°CQ3_K_M
✓ Measured
Qwen3 30B A3B Instruct 2507299.76 tok/s
166 W43°CQ4_K_M
✓ Measured
Qwen3-30B-A3B285.44 tok/s
160 W42°CQ4_K_M
✓ Measured
Qwen3-Coder 30B A3B296.22 tok/s
127 W37°CQ4_K_M
✓ Measured
Gemma 4 31B76.59 tok/s
324 W59°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 32B75.8 tok/s
210 W52°CQ4_K_M
✓ Measured
DeepSeek-R1-Distill-Qwen-32B-abliterated73.48 tok/s
344 W50°CQ4_K_M
✓ Measured
Olmo-3.1-32B-Think75.53 tok/s
344 W52°CQ4_K_M
✓ Measured
QwQ 32B75.82 tok/s
213 W51°CQ4_K_M
✓ Measured
Qwen2.5-32B73.5 tok/s
339 W51°CQ4_K_M
✓ Measured
Qwen2.5-Coder 32B75.81 tok/s
175 W52°CQ4_K_M
✓ Measured
Qwen2.5-Coder 32B (Q3_K_M)58.05 tok/s
351 W50°CQ3_K_M
✓ Measured
Qwen3 32B74.07 tok/s
19 GB peak218 W54°C0.34 tok/WQ4_K_M
✓ Measured
Dolphin 2.9.1 Yi 1.5 34B76.76 tok/s
199 W51°CQ4_K_M
✓ Measured
Ornith 1.5 35B A3B240.25 tok/s
165 W41°CQ4_K_M
✓ Measured
Ornith-1.0-35B221.65 tok/s
163 W42°CQ4_K_M
✓ Measured
Qwen-AgentWorld-35B-A3B219.1 tok/s
168 W43°CQ4_K_M
✓ Measured
Qwen3.6 35B A3B236.68 tok/s
163 W52°CQ4_K_M
✓ Measured
GLM-4.7-Flash186.05 tok/s
169 W42°CQ4_K_M
✓ Measured
Laguna-XS-2.1✕ Won't fit needs ~26 GBVRAM-gated at this precision✓ Measured
KAT-Coder-V2.5-Dev231.83 tok/s
166 W42°CQ4_K_M
✓ Measured
DeepSeek-R1-Distill-Llama-70B41.17 tok/s
380 W56°CQ4_K_M
✓ Measured
Hermes-4-70B41.2 tok/s
370 W55°CQ4_K_M
✓ Measured
Llama 3.3 70B41 tok/s
40.1 GB peak293 W57°C0.14 tok/WQ4_K_M
✓ Measured
Llama-3.3-70B-Instruct-abliterated41.15 tok/s
357 W56°CQ4_K_M
✓ Measured
Meta-Llama-3.1-70B41.19 tok/s
379 W52°CQ4_K_M
✓ Measured
Qwen2.5-72B40.78 tok/s
366 W56°CQ4_K_M
✓ Measured
Qwen3-Next-80B-A3B186.22 tok/s
148 W40°CQ4_K_M
✓ Measured
Qwen3-Next-80B-A3B-Thinking191.08 tok/s
148 W40°CQ4_K_M
✓ Measured
Qwen3-Next-80B-A3B-Thinking188.14 tok/s
151 W42°CQ4_K_M
✓ Measured
Qwen3-Coder-Next193.79 tok/s
149 W42°CQ4_K_M
✓ Measured
Qwen3-Coder-Next188.21 tok/s
152 W41°CQ4_K_M
✓ Measured
Qwen3-Coder-Next-abliterated190.32 tok/s
152 W41°CQ4_K_M
✓ Measured
gpt-oss-120b220.43 tok/s
154 W55°CMXFP4
✓ Measured

Image Generation images/min 15

SDXL Turbo618.27
FLUX.2 klein 4B108.33
Stable Diffusion 1.580.45
FLUX.2 klein 9B64.87
FLUX.1 Schnell58.36
Z-Image Turbo (1024px)39.42
Stable Diffusion XL34.58
Z-Image Turbo22.125
Playground v2.521.48
FLUX.1 dev9.064
Krea 2 Turbo6.85
Qwen-Image 2.14.59
WorkloadResultTelemetryData
Stable Diffusion 1.580.45 images/min
294 W31°C
✓ Measured
SDXL Turbo618.27 images/min
120 W25°C
✓ Measured
Stable Diffusion XL34.58 images/min
16.2 GB peak539 W56°C1.7 s/img
✓ Measured
Playground v2.521.48 images/min
543 W42°C
✓ Measured
FLUX.2 klein 4B108.33 images/min
512 W53°C
✓ Measured
Z-Image3.58 images/min
689 W53°C
✓ Measured
Z-Image Turbo (1024px)39.42 images/min
640 W45°C
✓ Measured
Z-Image Turbo22.13 images/min
26 GB peak661 W59°C2.7 s/img
✓ Measured
FLUX.1 Schnell58.36 images/min
614 W45°C
✓ Measured
FLUX.1 dev9.06 images/min
36.9 GB peak662 W63°C6.6 s/img
✓ Measured
FLUX.2 klein 9B64.87 images/min
649 W59°C
✓ Measured
Qwen-Image 2.14.59 images/min
690 W68°C
✓ Measured
Krea 2 Turbo6.85 images/min
638 W60°C
✓ Measured
Qwen-Image3 images/min
679 W63°C
✓ Measured
Qwen-Image 25123 images/min
669 W60°C
✓ Measured

Image Editing images/min 3

FLUX.1 Kontext dev4.264
Qwen-Image-Edit3.68
Qwen-Image-Edit 25091.65
WorkloadResultTelemetryData
FLUX.1 Kontext dev4.26 images/min
35.7 GB peak683 W65°C14.1 s/img
✓ Measured
Qwen-Image-Edit 25091.65 images/min
689 W65°C
✓ Measured
Qwen-Image-Edit3.68 images/min
60.4 GB peak683 W65°C16.3 s/img
✓ Measured

Background Removal images/min 1

WorkloadResultTelemetryData
BiRefNet1310.97 images/min
123 W33°C
✓ Measured

Upscaling images/min 1

WorkloadResultTelemetryData
Swin2SR 4x Upscaler25.48 images/min
203 W38°C
✓ Measured

Image to Video clips/min 7

LTX-Video (image to video)10.623
Stable Video Diffusion4.875
Wan 2.2 TI2V-5B (image to video)4.126
Stable Video Diffusion XT2.882
CogVideoX-5B I2V0.879
Wan 2.2 I2V A14B0.403
Wan 2.1 I2V 14B (480p)0.403
WorkloadResultTelemetryData
Stable Video Diffusion4.88 clips/min
615 W64°C
✓ Measured
LTX-Video (image to video)10.62 clips/min
595 W60°C
✓ Measured
Wan 2.2 TI2V-5B (image to video)4.13 clips/min
673 W64°C
✓ Measured
Stable Video Diffusion XT2.88 clips/min
619 W65°C
✓ Measured
CogVideoX-5B I2V0.88 clips/min
657 W67°C
✓ Measured
Wan 2.1 I2V 14B (480p)0.4 clips/min
692 W67°C
✓ Measured
Wan 2.2 I2V A14B0.4 clips/min
692 W81°C
✓ Measured

Video Generation frames/s 5

LTX-Video (distilled)17.47
Wan 2.2 5B (720p)1.33
CogVideoX-2B1.168
Mochi 1 Preview0.382
Wan 2.2 T2V A14B0.335
WorkloadResultTelemetryData
CogVideoX-2B1.17 frames/s
666 W74°C42 s/clip
✓ Measured
Mochi 1 Preview0.38 frames/s
683 W69°C128.1 s/clip
✓ Measured
LTX-Video (distilled)17.47 frames/s
60.3 GB peak554 W60°C5.6 s/clip
✓ Measured
Wan 2.2 5B (720p)1.33 frames/s
61 GB peak669 W67°C36.8 s/clip
✓ Measured
Wan 2.2 T2V A14B0.34 frames/s
689 W77°C146.1 s/clip
✓ Measured

Image to 3D assets/hour 4

TripoSR Image-to-3D2107.4
TripoSG Image-to-3D559.4
TRELLIS Image-to-3D167.2
TRELLIS.2 Image-to-3D87.9
WorkloadResultTelemetryData
TripoSR Image-to-3D2107.4 assets/hour
171 W39°C
✓ Measured
TripoSG Image-to-3D559.4 assets/hour
529 W70°C
✓ Measured
TRELLIS Image-to-3D167.2 assets/hour✓ Measured
TRELLIS.2 Image-to-3D87.9 assets/hour✓ Measured

Music Generation x realtime 4

ACE-Step 1.527.05
ACE-Step v1 3.5B14.83
DiffRhythm 25.63
MusicGen Small2.47
WorkloadResultTelemetryData
MusicGen Small2.47 x realtime
163 W37°C
✓ Measured
ACE-Step 1.527.05 x realtime
179 W32°C
✓ Measured
ACE-Step v1 3.5B14.83 x realtime
237 W34°C
✓ Measured
DiffRhythm 25.63 x realtime
170 W38°C
✓ Measured

Sound Effects x realtime 3

MOSS-SoundEffect v2.02.8
EzAudio XL1.58
MiDashengLM-Gen1.43
WorkloadResultTelemetryData
EzAudio XL1.58 x realtime
474 W41°C
✓ Measured
MOSS-SoundEffect v2.02.8 x realtime
577 W57°C
✓ Measured
MiDashengLM-Gen1.43 x realtime
230 W38°C
✓ Measured

Speech to Text x realtime 1

WorkloadResultTelemetryData
Whisper large-v3181.47 x realtime
138 W37°C
✓ Measured

Text to Speech x realtime 1

WorkloadResultTelemetryData
Kokoro TTS 82M222.02 x realtime
126 W37°C
✓ Measured

Depth Estimation images/min 2

WorkloadResultTelemetryData
Depth Anything V2 Small1081.54 images/min
116 W33°C
✓ Measured
Depth Anything V2 Large918.71 images/min
117 W34°C
✓ Measured

Segmentation images/min 2

WorkloadResultTelemetryData
SAM ViT-Base1431.27 images/min
130 W38°C
✓ Measured
SAM ViT-Huge320.64 images/min
285 W47°C
✓ Measured

Vision Language images/min 2

WorkloadResultTelemetryData
Florence-2 Base249.62 images/min
125 W35°C
✓ Measured
Florence-2 Large156.15 images/min
143 W36°C
✓ Measured

Fine-Tuning train tok/s 4

SmolLM2 1.7B LoRA18910.3
TinyLlama 1.1B LoRA17522
Qwen2.5 1.5B LoRA16034.9
Qwen2.5 7B LoRA8514.3
WorkloadResultTelemetryData
TinyLlama 1.1B LoRA17522 train tok/s
301 W46°C
✓ Measured
Qwen2.5 1.5B LoRA16034.9 train tok/s
335 W47°C
✓ Measured
SmolLM2 1.7B LoRA18910.3 train tok/s
398 W50°C
✓ Measured
Qwen2.5 7B LoRA8514.3 train tok/s
633 W56°C
✓ Measured

LLM Serving serve tok/s 4

TinyLlama 1.1B served8336.7
Qwen2.5 1.5B served6794.8
SmolLM2 1.7B served6509.9
Qwen2.5 7B served3601.1
WorkloadResultTelemetryData
TinyLlama 1.1B served8336.7 serve tok/s
197 W35°C
✓ Measured
Qwen2.5 1.5B served6794.8 serve tok/s
237 W37°C
✓ Measured
SmolLM2 1.7B served6509.9 serve tok/s
262 W39°C
✓ Measured
Qwen2.5 7B served3601.1 serve tok/s
424 W42°C
✓ Measured

Speculative Decoding x vs solo 1

WorkloadResultTelemetryData
Qwen2.5 1.5B + 0.5B draft0.73 x vs solo✓ Measured
How we measured this. Every result comes from our own pinned, reproducible AI suite, 12 workloads: the Qwen3-4B to Llama-70B LLM ladder (llama.cpp, Q4_K_M), SDXL / Z-Image / FLUX-dev generation, FLUX-Kontext / Qwen-Edit editing, and LTX / Wan video, run first-party on rented hardware with under 0.5% run-to-run variance. Peak VRAM, power draw, temperature and tokens-per-watt are captured per workload. “Won’t fit” rows are real data: where a model exceeds the card’s VRAM at the tested precision we record a hard gate rather than silently dropping to a smaller quant. Measured 2026-07-10 · harness 2.0.0.

NVIDIA H100 80GB HBM3 specifications

ArchitectureHopper
CUDA cores16,896
VRAM80GB HBM3
Memory bus5120-bit
Memory bandwidth3350 GB/s
Boost clock1,980 MHz
TDP700 W
ProcessTSMC 4N
InterfaceSXM5
Release date2022-09-20
Launch MSRP$30,000

Verdict, NVIDIA H100 80GB HBM3 on real AI workloads

NVIDIA H100 80GB HBM3 scores 62.6/100, #6 of 102. It ran all 12 workloads. Every figure here is our own measurement.

Relative performance: where the NVIDIA H100 80GB HBM3 lands

100% = this card, AI & Machine Learning headline metric (AI Score). #6 of 21 datacenter cards in this vertical.

GPURelative%AI Score
NVIDIA B200
125%78
NVIDIA B100
107%67
NVIDIA GH200 Grace Hopper
105%65.6
NVIDIA H200
104%65
NVIDIA H100 80GB HBM3
100%62.6
NVIDIA H800 80GB
100%62.6
NVIDIA H100 NVL
93%58.1
NVIDIA RTX PRO 6000 Blackwell Server Edition
80%49.9
NVIDIA H100 PCIe
80%49.8

← All AI & Machine Learning GPU rankings

The silicon

Transistors80,000 million
Die size814 mm²
Process node4 nm
Fabricated byTSMC
Transistor density98.3 million per mm²

Denser than 90% of the 746 cards we have silicon data for. Density is the clearest measure of what a process node bought: a card that gained it without growing the die got its speed from the fab rather than the architecture.

Silicon figures from Wikipedia (CC BY-SA 4.0). Benchmarks on this page are our own. Compare every chip.

What this card can build

Whole-job timings, composed from our measured per-model results on this card.

WorkflowTimeEnergyBasis
Depth pass on a batch5 s0.11 Whmeasured
Voiceovers from scripts17 s0.28 Whmeasured
Podcast episode pass34 s1.05 Whall 3 stages measured
Masking run1.6 min7.4 Whmeasured
Transcribe and subtitle videos1.7 min3.79 Whmeasured
24-frame storyboard1.8 min13.08 Whall 2 stages measured
60-second AI short film2.5 min17.46 Whall 3 stages measured
6-panel comic page3 min23.89 Whall 3 stages measured
Character sheet, 12 poses3.7 min33.26 Whall 2 stages measured
Animate a batch of images5.5 min54.4 Whmeasured
Caption a training dataset6.5 min15.22 Whmeasured
Full codebase review6.9 min26.8 Whmeasured
Short social clips8 min73.89 Whall 3 stages measured
Product catalogue cutout8.1 min26.87 Whall 2 stages measured
3D game asset kit8.6 min5.19 Whall 2 stages measured
Product photo shoot11.1 min117.19 Whall 2 stages measured
Long-form article batch11.4 min55.51 Whmeasured
Product shoot, start to finish12.9 min122.56 Whall 4 stages measured
Photos to 3D models18.9 min0.04 Whall 2 stages measured
Photo restoration batch23.8 min267 Whmeasured
Restore and enlarge photos27.8 min280.28 Whall 2 stages measured

Rent or buy?

This card is $30,000 to buy. The cheapest listed rate on Vast.ai is $2.003/hour, but that is the floor: we budget $2.404/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 12,481 GPU-hours. Below it you are paying for idle silicon.

How you would use itGPU-hours a yearRental cost a yearTime to break even
2 hours a day, hobby730$1,75517.1 years
8 hours a day, working on it2,920$7,0194.3 years
24/7, always-on agent8,760$21,0561.4 years

At hobby usage this card is very unlikely to pay for itself before it is superseded. Rent it. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.

Rental price

$2.690/hr+84.8% since 2026-08-14low $0.948 · high $2.690

Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.