NVIDIA A100 80GB SXM4, AI & Machine Learning Benchmarks & Specs

80GB · AI Score 33.1/100 · first-party measured on 12 AI workloads

33.1 AI Score ✓ Measured

Every number on this page is first-party: NVIDIA A100 80GB SXM4 was run on our pinned 12-workload AI suite on 2026-07-10, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) NVIDIA A100 80GB SXM4 delivers about 162.57 tokens/sec. Stepping up to Qwen3 32B it holds roughly 45.53 tok/s. The full Llama 3.3 70B still runs, at about 24.4 tok/s. For image generation, SDXL runs at 8.18 it/s, and FLUX.1-dev at 1.99 it/s. All 12 workloads fit in 80GB. There is no model in our suite this card has to turn down. NVIDIA A100 80GB SXM4 isn't a retail purchase for most people. It's rented by the hour. You can run this exact card on RunPod.

AI & Machine Learning benchmark results

Text Generation tok/s 123

gemma-3-270m617.58
LFM2.5-1.2B576.95
Qwen2.5-0.5B571.49
Qwen2.5-Coder-0.5B562.59
SmolLM2-135M554.44
Qwen1.5-0.5B540.32
Llama 3.2 1B525.32
Qwen3-0.6B432.66
Qwen3 0.6B428.39
LFM2.5-8B-A1B369.01
Qwen3-1.7B350.4
Qwen3 1.7B343.58
WorkloadResultTelemetryData
gemma-3-270m617.58 tok/s
103 W35°CQ4_K_M
✓ Measured
LFM2.5-1.2B576.95 tok/s
142 W44°CQ4_K_M
✓ Measured
Qwen2.5-0.5B571.49 tok/s
105 W35°CQ4_K_M
✓ Measured
Qwen2.5-Coder-0.5B562.59 tok/s
109 W35°CQ4_K_M
✓ Measured
SmolLM2-135M554.44 tok/s
104 W36°CQ4_K_M
✓ Measured
Qwen1.5-0.5B540.32 tok/s
110 W37°CQ4_K_M
✓ Measured
Llama 3.2 1B525.32 tok/s
99 W48°CQ4_K_M
✓ Measured
Qwen3-0.6B432.66 tok/s
105 W36°CQ4_K_M
✓ Measured
Qwen3 0.6B428.39 tok/s
96 W44°CQ4_K_M
✓ Measured
LFM2.5-8B-A1B369.01 tok/s
146 W39°CQ4_K_M
✓ Measured
Qwen3-1.7B350.4 tok/s
146 W39°CQ4_K_M
✓ Measured
Qwen3 1.7B343.58 tok/s
102 W48°CQ4_K_M
✓ Measured
Qwen2.5-1.5B335.29 tok/s
132 W37°CQ4_K_M
✓ Measured
Qwen2.5-Coder-1.5B333.87 tok/s
131 W41°CQ4_K_M
✓ Measured
Qwen2-1.5B333.79 tok/s
129 W39°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 1.5B327.66 tok/s
112 W48°CQ4_K_M
✓ Measured
gemma-3-1b324.08 tok/s
129 W37°CQ4_K_M
✓ Measured
Hermes-3-Llama-3.2-3B269.13 tok/s
179 W42°CQ4_K_M
✓ Measured
Llama-3.2-3B-Instruct-uncensored269.13 tok/s
138 W42°CQ4_K_M
✓ Measured
Llama 3.2 3B264.25 tok/s
131 W51°CQ4_K_M
✓ Measured
SmolLM3-3B254.6 tok/s
161 W41°CQ4_K_M
✓ Measured
SmolLM3 3B249.76 tok/s
122 W51°CQ4_K_M
✓ Measured
Qwen2.5-Coder-3B248.7 tok/s
170 W43°CQ4_K_M
✓ Measured
Qwen2.5-3B248.47 tok/s
151 W40°CQ4_K_M
✓ Measured
gemma-2-2b244.52 tok/s
149 W40°CQ4_K_M
✓ Measured
gemma-2-2b-it-abliterated244.41 tok/s
159 W41°CQ4_K_M
✓ Measured
Phi-4-mini243.38 tok/s
178 W43°CQ4_K_M
✓ Measured
AI21-Jamba-Reasoning-3B239.23 tok/s
176 W41°CQ4_K_M
✓ Measured
Phi-4 Mini 3.8B239.17 tok/s
126 W53°CQ4_K_M
✓ Measured
phi-2235.55 tok/s
161 W40°CQ4_K_M
✓ Measured
Phi-3.5-mini229.46 tok/s
209 W41°CQ4_K_M
✓ Measured
DeepSeek-Coder-V2-Lite213.91 tok/s
150 W41°CQ4_K_M
✓ Measured
gpt-oss-20b212.63 tok/s
154 W40°CQ4_K_M
✓ Measured
Nemotron-3-Nano-30B-A3B204.06 tok/s
132 W43°CQ4_K_M
✓ Measured
Qwen3-4B-Thinking-2507199.71 tok/s
184 W41°CQ4_K_M
✓ Measured
Qwen3-4B-Instruct-2507199.65 tok/s
152 W42°CQ4_K_M
✓ Measured
Qwen3-4B-Instruct-2507199.39 tok/s
155 W42°CQ4_K_M
✓ Measured
Qwen3 4B197 tok/s
3.2 GB peak145 W50°C1.36 tok/WQ4_K_M
✓ Measured
Qwen3-Coder 30B A3B182.26 tok/s
109 W51°CQ4_K_M
✓ Measured
Llama-2-7B180.09 tok/s
226 W47°CQ4_K_M
✓ Measured
Qwen3-30B-A3B178.14 tok/s
146 W39°CQ4_K_M
✓ Measured
Qwen3 30B A3B177.32 tok/s
101 W53°CQ4_K_M
✓ Measured
Gemma 3 4B176.94 tok/s
122 W50°CQ4_K_M
✓ Measured
Mistral-7B-Instruct-v0.1173.87 tok/s
224 W45°CQ4_K_M
✓ Measured
Mistral-7B-Instruct-v0.2173.8 tok/s
216 W45°CQ4_K_M
✓ Measured
Mistral-7B-Instruct-v0.3173.27 tok/s
219 W44°CQ4_K_M
✓ Measured
Mistral 7B v0.3171.92 tok/s
128 W55°CQ4_K_M
✓ Measured
Qwen2.5-7B166.92 tok/s
235 W45°CQ4_K_M
✓ Measured
Qwen2.5-Coder-7B-Instruct-abliterated166.84 tok/s
203 W44°CQ4_K_M
✓ Measured
Qwen2.5-Coder 7B164.6 tok/s
114 W55°CQ4_K_M
✓ Measured
L3-8B-Stheno-v3.2164.51 tok/s
200 W45°CQ4_K_M
✓ Measured
DarkIdol-Llama-3.1-8B-Instruct-1.2-Uncensored164.46 tok/s
230 W48°CQ4_K_M
✓ Measured
Meta-Llama-3.1-8B164.24 tok/s
222 W44°CQ4_K_M
✓ Measured
dolphin-2.9-llama3-8b164.09 tok/s
225 W45°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 7B162.96 tok/s
153 W50°CQ4_K_M
✓ Measured
Llama 3.1 8B162.57 tok/s
5.1 GB peak182 W55°C0.89 tok/WQ4_K_M
✓ Measured
DeepSeek-R1 Distill Llama 8B159.37 tok/s
119 W51°CQ4_K_M
✓ Measured
Dolphin 3.0 Llama 3.1 8B158.63 tok/s
174 W55°CQ4_K_M
✓ Measured
Dolphin X1 8B158.62 tok/s
156 W55°CQ4_K_M
✓ Measured
Dolphin X1 Trinity Nano 6B158.41 tok/s
131 W34°CQ4_K_M
✓ Measured
Qwen3-8B152.8 tok/s
223 W45°CQ4_K_M
✓ Measured
DeepSeek-R1-0528-Qwen3-8B152.76 tok/s
220 W44°CQ4_K_M
✓ Measured
Josiefied-Qwen3-8B-abliterated-v1152.29 tok/s
219 W44°CQ4_K_M
✓ Measured
Qwen3 8B149.96 tok/s
153 W48°CQ4_K_M
✓ Measured
Qwen3 30B A3B (Q3_K_M)146.72 tok/s
149 W40°CQ3_K_M
✓ Measured
KAT-Coder-V2.5-Dev146.38 tok/s
144 W39°CQ4_K_M
✓ Measured
Qwen-AgentWorld-35B-A3B139.69 tok/s
141 W38°CQ4_K_M
✓ Measured
Ornith-1.0-35B138.8 tok/s
135 W39°CQ4_K_M
✓ Measured
Ornith-1.0-9B134.67 tok/s
193 W45°CQ4_K_M
✓ Measured
GLM-4.7-Flash124.08 tok/s
146 W38°CQ4_K_M
✓ Measured
Qwen3-Next-80B-A3B-Thinking120.98 tok/s
128 W40°CQ4_K_M
✓ Measured
Qwen3-Coder-Next120.71 tok/s
130 W39°CQ4_K_M
✓ Measured
Qwen3-Coder-Next-abliterated119.56 tok/s
129 W40°CQ4_K_M
✓ Measured
Qwen3-Next-80B-A3B-Thinking118.93 tok/s
127 W40°CQ4_K_M
✓ Measured
Qwen3-Coder-Next118.62 tok/s
127 W39°CQ4_K_M
✓ Measured
Qwen3-Next-80B-A3B117.44 tok/s
126 W40°CQ4_K_M
✓ Measured
GLM-4.7-Flash-REAP-23B-A3B112.75 tok/s
144 W38°CQ4_K_M
✓ Measured
Mistral-Nemo-Instruct-2407111.07 tok/s
230 W46°CQ4_K_M
✓ Measured
NemoMix-Unleashed-12B110.99 tok/s
236 W46°CQ4_K_M
✓ Measured
gemma-2-9b108.68 tok/s
221 W44°CQ4_K_M
✓ Measured
Phi-4 14B95.22 tok/s
173 W57°CQ4_K_M
✓ Measured
Hermes-4-14B94.48 tok/s
244 W47°CQ4_K_M
✓ Measured
Qwen3-14B94.35 tok/s
245 W46°CQ4_K_M
✓ Measured
Gemma 4 12B92.43 tok/s
167 W56°CQ4_K_M
✓ Measured
Gemma 3 12B91.45 tok/s
159 W53°CQ4_K_M
✓ Measured
Qwen3 14B90.64 tok/s
163 W54°CQ4_K_M
✓ Measured
Qwen2.5-14B90.46 tok/s
247 W46°CQ4_K_M
✓ Measured
EVA-Qwen2.5-14B-v0.290.35 tok/s
242 W47°CQ4_K_M
✓ Measured
Qwen2.5-Coder-14B-Instruct-abliterated90.28 tok/s
231 W46°CQ4_K_M
✓ Measured
Qwen2.5-Coder 14B89.45 tok/s
8.8 GB peak175 W56°C0.51 tok/WQ4_K_M
✓ Measured
Uncensored89.43 tok/s
211 W46°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 14B87 tok/s
149 W57°CQ4_K_M
✓ Measured
Phi-4 14B (Q3_K_M)82.68 tok/s
251 W48°CQ3_K_M
✓ Measured
StarCoder2 15B79.22 tok/s
178 W54°CQ4_K_M
✓ Measured
Gemma 3 12B (Q3_K_M)71.88 tok/s
234 W45°CQ3_K_M
✓ Measured
DeepSeek-R1 Distill 14B (Q3_K_M)67.13 tok/s
235 W46°CQ3_K_M
✓ Measured
Cydonia-24B-v4.365.82 tok/s
255 W48°CQ4_K_M
✓ Measured
Dolphin-Mistral-24B-Venice-Edition65.77 tok/s
263 W48°CQ4_K_M
✓ Measured
Codestral 22B63.26 tok/s
174 W53°CQ4_K_M
✓ Measured
Dolphin Mistral 24B Venice62.72 tok/s
161 W50°CQ4_K_M
✓ Measured
Devstral Small 24B61.63 tok/s
186 W51°CQ4_K_M
✓ Measured
Mistral Small 24B61.38 tok/s
154 W58°CQ4_K_M
✓ Measured
Dolphin 3.0 R1 Mistral 24B61.17 tok/s
187 W56°CQ4_K_M
✓ Measured
Codestral 22B (Q3_K_M)49.52 tok/s
273 W49°CQ3_K_M
✓ Measured
Gemma 3 27B48.43 tok/s
169 W60°CQ4_K_M
✓ Measured
Mistral Small 24B (Q3_K_M)46.63 tok/s
259 W48°CQ3_K_M
✓ Measured
Olmo-3.1-32B-Think46.49 tok/s
261 W50°CQ4_K_M
✓ Measured
Qwen2.5-32B45.6 tok/s
266 W49°CQ4_K_M
✓ Measured
Qwen3 32B45.53 tok/s
18.9 GB peak158 W59°C0.29 tok/WQ4_K_M
✓ Measured
DeepSeek-R1-Distill-Qwen-32B-abliterated43.48 tok/s
233 W47°CQ4_K_M
✓ Measured
Dolphin 2.9.1 Yi 1.5 34B43.47 tok/s
164 W54°CQ4_K_M
✓ Measured
Qwen2.5-Coder 32B43.36 tok/s
174 W58°CQ4_K_M
✓ Measured
QwQ 32B43.27 tok/s
160 W59°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 32B43.11 tok/s
186 W61°CQ4_K_M
✓ Measured
Qwen2.5-Coder 32B (Q3_K_M)31.99 tok/s
249 W50°CQ3_K_M
✓ Measured
Hermes-4-70B24.47 tok/s
267 W52°CQ4_K_M
✓ Measured
Llama 3.3 70B24.4 tok/s
40 GB peak175 W62°C0.14 tok/WQ4_K_M
✓ Measured
Qwen2.5-72B23.95 tok/s
268 W52°CQ4_K_M
✓ Measured
DeepSeek-R1-Distill-Llama-70B22.99 tok/s
229 W48°CQ4_K_M
✓ Measured
Llama-3.3-70B-Instruct-abliterated22.92 tok/s
231 W50°CQ4_K_M
✓ Measured
Meta-Llama-3.1-70B22.92 tok/s
235 W50°CQ4_K_M
✓ Measured
Laguna-XS-2.1✕ Won't fit needs ~26 GBVRAM-gated at this precision✓ Measured
Nanbeige4.2-3B✕ Won't fit needs ~4 GBVRAM-gated at this precision✓ Measured

Image Generation images/min 6

FLUX.1 Schnell28.39
Z-Image Turbo (1024px)18.65
Stable Diffusion XL16.36
Z-Image Turbo9.975
FLUX.1 dev4.264
Krea 2 Turbo3.35
WorkloadResultTelemetryData
FLUX.1 Schnell28.39 images/min
394 W65°C
✓ Measured
Z-Image Turbo (1024px)18.65 images/min
396 W68°C
✓ Measured
Stable Diffusion XL16.36 images/min
14.7 GB peak370 W61°C3.7 s/img
✓ Measured
Z-Image Turbo9.98 images/min
25.8 GB peak382 W67°C6 s/img
✓ Measured
FLUX.1 dev4.26 images/min
36.7 GB peak396 W71°C14 s/img
✓ Measured
Krea 2 Turbo3.35 images/min
346 W63°C
✓ Measured

Fine-Tuning train tok/s 4

SmolLM2 1.7B LoRA9750.8
TinyLlama 1.1B LoRA9553.3
Qwen2.5 1.5B LoRA8252.8
Qwen2.5 7B LoRA3900.6
WorkloadResultTelemetryData
SmolLM2 1.7B LoRA9750.8 train tok/s
326 W44°C
✓ Measured
TinyLlama 1.1B LoRA9553.3 train tok/s
259 W39°C
✓ Measured
Qwen2.5 1.5B LoRA8252.8 train tok/s
260 W40°C
✓ Measured
Qwen2.5 7B LoRA3900.6 train tok/s
394 W49°C
✓ Measured

LLM Serving serve tok/s 4

TinyLlama 1.1B served5701.5
Qwen2.5 1.5B served4564.7
SmolLM2 1.7B served4351.4
Qwen2.5 7B served2187.2
WorkloadResultTelemetryData
TinyLlama 1.1B served5701.5 serve tok/s
179 W39°C
✓ Measured
Qwen2.5 1.5B served4564.7 serve tok/s
192 W41°C
✓ Measured
SmolLM2 1.7B served4351.4 serve tok/s
205 W42°C
✓ Measured
Qwen2.5 7B served2187.2 serve tok/s
276 W45°C
✓ Measured

Image to Video clips/min 3

Wan 2.2 TI2V-5B (image to video)2.156
Stable Video Diffusion XT1.536
CogVideoX-5B I2V0.468
WorkloadResultTelemetryData
Wan 2.2 TI2V-5B (image to video)2.16 clips/min
455 W66°C
✓ Measured
Stable Video Diffusion XT1.54 clips/min
431 W67°C
✓ Measured
CogVideoX-5B I2V0.47 clips/min
458 W70°C
✓ Measured

Image Editing images/min 2

WorkloadResultTelemetryData
FLUX.1 Kontext dev1.99 images/min
35.5 GB peak395 W73°C30 s/img
✓ Measured
Qwen-Image-Edit1.64 images/min
60.2 GB peak393 W74°C36.4 s/img
✓ Measured

Video Generation frames/s 2

WorkloadResultTelemetryData
LTX-Video (distilled)8.95 frames/s
60.1 GB peak379 W69°C10.8 s/clip
✓ Measured
Wan 2.2 5B (720p)0.66 frames/s
60.3 GB peak392 W74°C74.2 s/clip
✓ Measured

Image to 3D assets/hour 2

WorkloadResultTelemetryData
TRELLIS Image-to-3D488.7 assets/hour
225 W54°C
✓ Measured
TRELLIS.2 Image-to-3D (1536³ max quality)32.7 assets/hour
292 W77°C
✓ Measured

Depth Estimation images/min 2

WorkloadResultTelemetryData
Depth Anything V2 Small925.15 images/min
81 W50°C
✓ Measured
Depth Anything V2 Large876.29 images/min
82 W50°C
✓ Measured

Segmentation images/min 2

WorkloadResultTelemetryData
SAM ViT-Base744.79 images/min
80 W51°C
✓ Measured
SAM ViT-Huge145.7 images/min
303 W58°C
✓ Measured

Vision Language images/min 2

WorkloadResultTelemetryData
Florence-2 Base126.4 images/min
79 W34°C
✓ Measured
Florence-2 Large70.94 images/min
82 W35°C
✓ Measured

Background Removal images/min 1

WorkloadResultTelemetryData
BiRefNet916.67 images/min
83 W50°C
✓ Measured

Upscaling images/min 1

WorkloadResultTelemetryData
Swin2SR 4x Upscaler28.72 images/min
184 W56°C
✓ Measured

Text to Speech x realtime 1

WorkloadResultTelemetryData
Kokoro TTS 82M110.17 x realtime
72 W40°C
✓ Measured

Music Generation x realtime 1

WorkloadResultTelemetryData
MusicGen Small0.74 x realtime
84 W40°C
✓ Measured

Speech to Text x realtime 1

WorkloadResultTelemetryData
Whisper large-v373.83 x realtime
151 W45°C
✓ Measured
How we measured this. Every result comes from our own pinned, reproducible AI suite, 12 workloads: the Qwen3-4B to Llama-70B LLM ladder (llama.cpp, Q4_K_M), SDXL / Z-Image / FLUX-dev generation, FLUX-Kontext / Qwen-Edit editing, and LTX / Wan video, run first-party on rented hardware with under 0.5% run-to-run variance. Peak VRAM, power draw, temperature and tokens-per-watt are captured per workload. “Won’t fit” rows are real data: where a model exceeds the card’s VRAM at the tested precision we record a hard gate rather than silently dropping to a smaller quant. Measured 2026-07-10 · harness 2.0.0.

NVIDIA A100 80GB SXM4 specifications

ArchitectureAmpere
CUDA cores6,912
VRAM80GB HBM2e
Memory bus5120-bit
Memory bandwidth2039 GB/s
Boost clock1,410 MHz
TDP400 W
ProcessTSMC 7nm
InterfaceSXM4
Release date2020-11-16
Launch MSRP$17,000

Verdict, NVIDIA A100 80GB SXM4 on real AI workloads

NVIDIA A100 80GB SXM4 scores 33.1/100, #13 of 102. It ran all 12 workloads. Every figure here is our own measurement.

Relative performance: where the NVIDIA A100 80GB SXM4 lands

100% = this card, AI & Machine Learning headline metric (AI Score). #11 of 21 desktop cards in this vertical.

GPURelative%AI Score
NVIDIA H100 80GB HBM3
189%62.6
NVIDIA H800 80GB
189%62.6
NVIDIA RTX PRO 6000 Blackwell Server Edition
151%49.9
NVIDIA H100 PCIe
140%46.4
NVIDIA A100 80GB SXM4
100%33.1
NVIDIA A800 80GB
100%33.1
NVIDIA A100 80GB PCIe
96%31.7
NVIDIA L40S
84%27.7
NVIDIA L40
59%19.5

← All AI & Machine Learning GPU rankings

What this card can build

Whole-job timings, composed from our measured per-model results on this card.

WorkflowTimeEnergyBasis
50-image depth pass5 s0.08 Whmeasured
30-minute podcast pass67 s1.81 Whall 3 stages measured
500-image masking run3.5 min17.35 Whmeasured
24-frame storyboard3.7 min16.68 Whall 2 stages measured
20-asset 3D game kit4 min16.75 Whall 2 stages measured
60-second AI short film4.7 min23.61 Whall 3 stages measured
60-second AI short film, narrated4.9 min23.62 Whall 4 stages measured
6-panel comic page6 min29.8 Whall 3 stages measured
200-product catalogue cutout7.3 min21.66 Whall 2 stages measured
Character sheet, 12 poses7.6 min41.23 Whall 2 stages measured
Full codebase review11.2 min32.59 Whmeasured
10 short social clips15.1 min87.69 Whall 3 stages measured
20 long-form articles19.1 min55.72 Whmeasured
40-product photo shoot23.5 min147.33 Whall 2 stages measured
40-product shoot, start to finish25 min151.66 Whall 4 stages measured
100-photo restoration batch50.8 min330.66 Whmeasured
100-photo restore and enlarge54.3 min341.34 Whall 2 stages measured

Rent or buy?

This card is $17,000 to buy. The cheapest listed rate on RunPod is $1.390/hour, but that is the floor: we budget $1.668/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 10,192 GPU-hours. Below it you are paying for idle silicon.

How you would use itGPU-hours a yearRental cost a yearTime to break even
2 hours a day, hobby730$1,21814.0 years
8 hours a day, working on it2,920$4,8713.5 years
24/7, always-on agent8,760$14,6121.2 years

At hobby usage this card is very unlikely to pay for itself before it is superseded. Rent it. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.

Rental price

$1.390/hr+0.0% since 2026-08-14low $1.390 · high $1.390

Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.