NVIDIA A100 40GB SXM4, AI & Machine Learning Benchmarks & Specs

40GB · AI Score 17.0/100 · anchored estimate vs 51 measured cards

17 AI Score Includes estimates

We have not run NVIDIA A100 40GB SXM4 on our bench. These figures are anchored estimates, interpolated per workload against the 51 GPUs we did measure (confidence: high (sibling silicon)). On Llama 3.1 8B (Q4_K_M) NVIDIA A100 40GB SXM4 should deliver about 124 tokens/sec. Stepping up to Qwen3 32B it should hold roughly 34.7 tok/s. Llama 3.3 70B does not fit. It needs roughly 42GB and this card has 40GB. For image generation, SDXL should run near 8.18 it/s, and FLUX.1-dev at 1.99 it/s. 2 of the 12 workloads won't fit on 40GB at the tested precision, Llama 3.3 70B, Qwen-Image-Edit. We publish those as hard gates rather than quietly dropping to a smaller quant.

AI & Machine Learning benchmark results

Text Generation tok/s 123

gemma-3-270m574.24
LFM2.5-1.2B572.69
SmolLM2-135M544.06
Qwen2.5-Coder-0.5B542.36
Llama 3.2 1B532.08
Qwen2.5-0.5B517.09
Qwen1.5-0.5B497.73
Qwen3 0.6B417.88
Qwen3-0.6B408.39
LFM2.5-8B-A1B351.47
Qwen3-1.7B341.85
Qwen3 1.7B340.76
WorkloadResultTelemetryData
gemma-3-270m574.24 tok/s
74 W36°CQ4_K_M
✓ Measured
LFM2.5-1.2B572.69 tok/s
94 W42°CQ4_K_M
✓ Measured
SmolLM2-135M544.06 tok/s
69 W38°CQ4_K_M
✓ Measured
Qwen2.5-Coder-0.5B542.36 tok/s
75 W37°CQ4_K_M
✓ Measured
Llama 3.2 1B532.08 tok/s
113 W41°CQ4_K_M
✓ Measured
Qwen2.5-0.5B517.09 tok/s
69 W36°CQ4_K_M
✓ Measured
Qwen1.5-0.5B497.73 tok/s
71 W36°CQ4_K_M
✓ Measured
Qwen3 0.6B417.88 tok/s
70 W36°CQ4_K_M
✓ Measured
Qwen3-0.6B408.39 tok/s
76 W38°CQ4_K_M
✓ Measured
LFM2.5-8B-A1B351.47 tok/s
93 W42°CQ4_K_M
✓ Measured
Qwen3-1.7B341.85 tok/s
103 W41°CQ4_K_M
✓ Measured
Qwen3 1.7B340.76 tok/s
112 W40°CQ4_K_M
✓ Measured
Qwen2.5-Coder-1.5B322.96 tok/s
95 W39°CQ4_K_M
✓ Measured
Qwen2-1.5B322.09 tok/s
93 W41°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 1.5B320.78 tok/s
93 W41°CQ4_K_M
✓ Measured
Qwen2.5-1.5B315.59 tok/s
100 W40°CQ4_K_M
✓ Measured
gemma-3-1b310.69 tok/s
88 W38°CQ4_K_M
✓ Measured
Llama 3.2 3B261.24 tok/s
118 W44°CQ4_K_M
✓ Measured
Hermes-3-Llama-3.2-3B261 tok/s
126 W44°CQ4_K_M
✓ Measured
Llama-3.2-3B-Instruct-uncensored260.83 tok/s
115 W43°CQ4_K_M
✓ Measured
SmolLM3-3B247.66 tok/s
131 W43°CQ4_K_M
✓ Measured
SmolLM3 3B246.25 tok/s
129 W43°CQ4_K_M
✓ Measured
Qwen2.5-Coder-3B241.51 tok/s
129 W43°CQ4_K_M
✓ Measured
Qwen2.5-3B240.48 tok/s
127 W43°CQ4_K_M
✓ Measured
Phi-4-mini237.38 tok/s
125 W46°CQ4_K_M
✓ Measured
Phi-4 Mini 3.8B236.71 tok/s
139 W46°CQ4_K_M
✓ Measured
gemma-2-2b234.68 tok/s
109 W43°CQ4_K_M
✓ Measured
gemma-2-2b-it-abliterated232.27 tok/s
121 W43°CQ4_K_M
✓ Measured
AI21-Jamba-Reasoning-3B231.15 tok/s
110 W43°CQ4_K_M
✓ Measured
phi-2228.1 tok/s
131 W43°CQ4_K_M
✓ Measured
Phi-3.5-mini222.53 tok/s
132 W44°CQ4_K_M
✓ Measured
DeepSeek-Coder-V2-Lite205.42 tok/s
120 W43°CQ4_K_M
✓ Measured
gpt-oss-20b205.1 tok/s
121 W45°CQ4_K_M
✓ Measured
Qwen3-4B-Instruct-2507193.39 tok/s
140 W45°CQ4_K_M
✓ Measured
Qwen3-4B193.14 tok/s
137 W44°CQ4_K_M
✓ Measured
Qwen3-4B-Instruct-2507192.89 tok/s
133 W45°CQ4_K_M
✓ Measured
Nemotron-3-Nano-30B-A3B192.44 tok/s
103 W43°CQ4_K_M
✓ Measured
Qwen3-4B-Thinking-2507189.89 tok/s
134 W44°CQ4_K_M
✓ Measured
Gemma 3 4B174.14 tok/s
138 W43°CQ4_K_M
✓ Measured
Qwen3-Coder 30B A3B173.47 tok/s
92 W42°CQ4_K_M
✓ Measured
Llama-2-7B172.53 tok/s
195 W49°CQ4_K_M
✓ Measured
Qwen3 30B A3B169.1 tok/s
92 W42°CQ4_K_M
✓ Measured
Qwen3-30B-A3B167.1 tok/s
106 W42°CQ4_K_M
✓ Measured
Mistral-7B-Instruct-v0.3167.06 tok/s
167 W47°CQ4_K_M
✓ Measured
Mistral-7B-Instruct-v0.2166.8 tok/s
165 W47°CQ4_K_M
✓ Measured
Mistral-7B-Instruct-v0.1166.39 tok/s
174 W48°CQ4_K_M
✓ Measured
Mistral 7B v0.3165.9 tok/s
182 W48°CQ4_K_M
✓ Measured
Qwen2.5-Coder 7B159.9 tok/s
176 W47°CQ4_K_M
✓ Measured
Qwen2.5-7B159.81 tok/s
172 W48°CQ4_K_M
✓ Measured
Qwen2.5-Coder-7B-Instruct-abliterated159.76 tok/s
151 W47°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 7B159.09 tok/s
164 W47°CQ4_K_M
✓ Measured
Llama-3.1-8B157.26 tok/s
170 W47°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill Llama 8B157.16 tok/s
166 W48°CQ4_K_M
✓ Measured
Meta-Llama-3.1-8B156.68 tok/s
170 W48°CQ4_K_M
✓ Measured
DarkIdol-Llama-3.1-8B-Instruct-1.2-Uncensored156.62 tok/s
177 W47°CQ4_K_M
✓ Measured
Dolphin 3.0 Llama 3.1 8B156.55 tok/s
138 W49°CQ4_K_M
✓ Measured
L3-8B-Stheno-v3.2156.55 tok/s
163 W49°CQ4_K_M
✓ Measured
Dolphin X1 8B156.3 tok/s
153 W49°CQ4_K_M
✓ Measured
dolphin-2.9-llama3-8b153.34 tok/s
140 W48°CQ4_K_M
✓ Measured
Dolphin X1 Trinity Nano 6B153.27 tok/s
91 W39°CQ4_K_M
✓ Measured
Josiefied-Qwen3-8B-abliterated-v1146.8 tok/s
168 W47°CQ4_K_M
✓ Measured
Qwen3 8B146.21 tok/s
164 W46°CQ4_K_M
✓ Measured
Qwen3-8B145.96 tok/s
165 W47°CQ4_K_M
✓ Measured
DeepSeek-R1-0528-Qwen3-8B145.84 tok/s
142 W47°CQ4_K_M
✓ Measured
Qwen3 30B A3B (Q3_K_M)142.78 tok/s
115 W43°CQ3_K_M
✓ Measured
KAT-Coder-V2.5-Dev137.31 tok/s
107 W42°CQ4_K_M
✓ Measured
Ornith-1.0-35B131.77 tok/s
104 W41°CQ4_K_M
✓ Measured
Qwen-AgentWorld-35B-A3B131.16 tok/s
97 W43°CQ4_K_M
✓ Measured
Ornith-1.0-9B126.79 tok/s
170 W48°CQ4_K_M
✓ Measured
GLM-4.7-Flash119.1 tok/s
98 W41°CQ4_K_M
✓ Measured
GLM-4.7-Flash-REAP-23B-A3B109.1 tok/s
103 W42°CQ4_K_M
✓ Measured
Mistral-Nemo-Instruct-2407106.33 tok/s
171 W49°CQ4_K_M
✓ Measured
NemoMix-Unleashed-12B105.05 tok/s
161 W49°CQ4_K_M
✓ Measured
gemma-2-9b101.97 tok/s
175 W47°CQ4_K_M
✓ Measured
Phi-4 14B98.65 tok/s
200 W52°CQ4_K_M
✓ Measured
Qwen3-14B90.07 tok/s
188 W50°CQ4_K_M
✓ Measured
Qwen3 14B90.01 tok/s
206 W50°CQ4_K_M
✓ Measured
Hermes-4-14B89.97 tok/s
166 W50°CQ4_K_M
✓ Measured
Gemma 3 12B88.97 tok/s
175 W48°CQ4_K_M
✓ Measured
Gemma 4 12B87.14 tok/s
162 W47°CQ4_K_M
✓ Measured
EVA-Qwen2.5-14B-v0.286.48 tok/s
187 W50°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 14B86.39 tok/s
167 W50°CQ4_K_M
✓ Measured
Qwen2.5-Coder-14B-Instruct-abliterated86.34 tok/s
183 W49°CQ4_K_M
✓ Measured
Qwen2.5-Coder-14B86.2 tok/s
182 W49°CQ4_K_M
✓ Measured
Qwen2.5-14B86.07 tok/s
193 W49°CQ4_K_M
✓ Measured
Uncensored85.82 tok/s
188 W50°CQ4_K_M
✓ Measured
Phi-4 14B (Q3_K_M)80.71 tok/s
182 W50°CQ3_K_M
✓ Measured
StarCoder2 15B80.25 tok/s
217 W52°CQ4_K_M
✓ Measured
Gemma 3 12B (Q3_K_M)69.7 tok/s
183 W48°CQ3_K_M
✓ Measured
DeepSeek-R1 Distill 14B (Q3_K_M)65.33 tok/s
187 W49°CQ3_K_M
✓ Measured
Codestral 22B63.74 tok/s
203 W52°CQ4_K_M
✓ Measured
Dolphin-Mistral-24B-Venice-Edition62.56 tok/s
212 W52°CQ4_K_M
✓ Measured
Dolphin Mistral 24B Venice62.46 tok/s
214 W53°CQ4_K_M
✓ Measured
Mistral Small 24B62.42 tok/s
204 W53°CQ4_K_M
✓ Measured
Cydonia-24B-v4.362.41 tok/s
195 W53°CQ4_K_M
✓ Measured
Devstral Small 24B62.4 tok/s
196 W52°CQ4_K_M
✓ Measured
Dolphin 3.0 R1 Mistral 24B62.4 tok/s
204 W52°CQ4_K_M
✓ Measured
Codestral 22B (Q3_K_M)48.37 tok/s
210 W52°CQ3_K_M
✓ Measured
Gemma 3 27B47.47 tok/s
200 W52°CQ4_K_M
✓ Measured
Mistral Small 24B (Q3_K_M)45.34 tok/s
200 W52°CQ3_K_M
✓ Measured
Qwen3-32B43.8 tok/s
235 W54°CQ4_K_M
✓ Measured
Olmo-3.1-32B-Think43.67 tok/s
212 W53°CQ4_K_M
✓ Measured
Qwen2.5-32B43.23 tok/s
206 W53°CQ4_K_M
✓ Measured
Qwen2.5-Coder 32B43.14 tok/s
208 W53°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 32B43.13 tok/s
212 W54°CQ4_K_M
✓ Measured
DeepSeek-R1-Distill-Qwen-32B-abliterated43.07 tok/s
192 W53°CQ4_K_M
✓ Measured
QwQ 32B43.02 tok/s
220 W53°CQ4_K_M
✓ Measured
Dolphin 2.9.1 Yi 1.5 34B42.97 tok/s
233 W54°CQ4_K_M
✓ Measured
Qwen2.5-Coder 32B (Q3_K_M)31.04 tok/s
210 W53°CQ3_K_M
✓ Measured
Llama-3.3-70B✕ Won't fit needs ~54 GBVRAM-gated at this precision✓ Measured
Qwen3-Coder-Next-abliterated✕ Won't fit needs ~61 GBVRAM-gated at this precision✓ Measured
Laguna-XS-2.1✕ Won't fit needs ~26 GBVRAM-gated at this precision✓ Measured
Nanbeige4.2-3B✕ Won't fit needs ~4 GBVRAM-gated at this precision✓ Measured
Qwen3-Coder-Next✕ Won't fit needs ~61 GBVRAM-gated at this precision✓ Measured
Qwen3-Next-80B-A3B-Thinking✕ Won't fit needs ~61 GBVRAM-gated at this precision✓ Measured
Qwen2.5-72B✕ Won't fit needs ~60 GBVRAM-gated at this precision✓ Measured
Qwen3-Coder-Next✕ Won't fit needs ~61 GBVRAM-gated at this precision✓ Measured
Qwen3-Next-80B-A3B-Thinking✕ Won't fit needs ~61 GBVRAM-gated at this precision✓ Measured
Qwen3-Next-80B-A3B✕ Won't fit needs ~61 GBVRAM-gated at this precision✓ Measured
DeepSeek-R1-Distill-Llama-70B✕ Won't fit needs ~54 GBVRAM-gated at this precision✓ Measured
Hermes-4-70B✕ Won't fit needs ~54 GBVRAM-gated at this precision✓ Measured
Llama-3.3-70B-Instruct-abliterated✕ Won't fit needs ~54 GBVRAM-gated at this precision✓ Measured
Meta-Llama-3.1-70B✕ Won't fit needs ~54 GBVRAM-gated at this precision✓ Measured

Image Generation images/min 5

FLUX.1 Schnell28.44
Z-Image Turbo (1024px)18.7
Stable Diffusion XL18.576
Z-Image Turbo9.975
FLUX.1 dev4.264
WorkloadResultTelemetryData
FLUX.1 Schnell28.44 images/min
384 W56°C
✓ Measured
Z-Image Turbo (1024px)18.7 images/min
382 W61°C
✓ Measured
Stable Diffusion XL18.58 images/min
10.3 GB peak375 W61°C3.2 s/img
✓ Measured
Z-Image Turbo9.98 images/minestimatedEst.
FLUX.1 dev4.26 images/minestimatedEst.

Fine-Tuning train tok/s 4

TinyLlama 1.1B LoRA7471
SmolLM2 1.7B LoRA6938.8
Qwen2.5 1.5B LoRA6349.7
Qwen2.5 7B LoRA3670.8
WorkloadResultTelemetryData
TinyLlama 1.1B LoRA7471 train tok/s
189 W57°C
✓ Measured
SmolLM2 1.7B LoRA6938.8 train tok/s
254 W60°C
✓ Measured
Qwen2.5 1.5B LoRA6349.7 train tok/s
210 W58°C
✓ Measured
Qwen2.5 7B LoRA3670.8 train tok/s
381 W67°C
✓ Measured

LLM Serving serve tok/s 4

Qwen2.5 1.5B served4111.1
SmolLM2 1.7B served3865.2
TinyLlama 1.1B served3221.3
Qwen2.5 7B served1769.6
WorkloadResultTelemetryData
Qwen2.5 1.5B served4111.1 serve tok/s
178 W45°C
✓ Measured
SmolLM2 1.7B served3865.2 serve tok/s
187 W45°C
✓ Measured
TinyLlama 1.1B served3221.3 serve tok/s
153 W40°C
✓ Measured
Qwen2.5 7B served1769.6 serve tok/s
225 W49°C
✓ Measured

Image Editing images/min 2

WorkloadResultTelemetryData
FLUX.1 Kontext dev1.99 images/minestimatedEst.
Qwen-Image-Edit✕ Won't fit VRAM-gated at this precisionEst.

Video Generation frames/s 2

WorkloadResultTelemetryData
LTX-Video (distilled)8.95 frames/sestimatedEst.
Wan 2.2 5B (720p)0.66 frames/sestimatedEst.

Depth Estimation images/min 2

WorkloadResultTelemetryData
Depth Anything V2 Small613.42 images/min
66 W35°C
✓ Measured
Depth Anything V2 Large536.74 images/min
74 W35°C
✓ Measured

Segmentation images/min 2

WorkloadResultTelemetryData
SAM ViT-Base696.77 images/min
64 W37°C
✓ Measured
SAM ViT-Huge146.81 images/min
318 W52°C
✓ Measured

Background Removal images/min 1

WorkloadResultTelemetryData
BiRefNet611.56 images/min
76 W35°C
✓ Measured

Upscaling images/min 1

WorkloadResultTelemetryData
Swin2SR 4x Upscaler17.25 images/min
116 W44°C
✓ Measured

Text to Speech x realtime 1

WorkloadResultTelemetryData
Kokoro TTS 82M142.03 x realtime
60 W32°C
✓ Measured

Music Generation x realtime 1

WorkloadResultTelemetryData
MusicGen Small1.08 x realtime
94 W34°C
✓ Measured

Speech to Text x realtime 1

WorkloadResultTelemetryData
Whisper large-v3102.47 x realtime
190 W36°C
✓ Measured

Speculative Decoding x vs solo 1

WorkloadResultTelemetryData
Qwen2.5 1.5B + 0.5B draft0.69 x vs solo✓ Measured
How this estimate is derived. This card hasn’t been through our bench yet, so its numbers are anchored estimates, interpolated from the 51 first-party measured cards (Sibling-anchored to measured A100 80GB: LLM ×0.763 (bandwidth ratio), diffusion ×1.0 (identical compute); VRAM gates at 40GB). The VRAM “won’t fit” gates are exact, since they’re pure capacity limits. Confidence: high (sibling silicon). Estimates are replaced with measured data as more silicon goes through the bench. Full methodology →

NVIDIA A100 40GB SXM4 specifications

ArchitectureAmpere
CUDA cores6,912
VRAM40GB HBM2
Memory bus5120-bit
Memory bandwidth1555 GB/s
Boost clock1,410 MHz
TDP400 W
ProcessTSMC 7nm
InterfaceSXM4
Release date2020-05-14
Launch MSRP$12,000

Verdict, capable, but 40GB sets the ceiling

NVIDIA A100 40GB SXM4 scores 17.0/100, #23 of 102. It ran 10 of 12; 2 exceeded its 40GB. Figures are anchored estimates, not measurements, we flag that on every row.

Relative performance: where the NVIDIA A100 40GB SXM4 lands

100% = this card, AI & Machine Learning headline metric (AI Score). #17 of 21 desktop cards in this vertical.

GPURelative%AI Score
NVIDIA A100 80GB PCIe
186%31.7
NVIDIA L40S
163%27.7
NVIDIA L40
115%19.5
NVIDIA A40
104%17.7
NVIDIA A100 40GB SXM4
100%17
NVIDIA A100 40GB PCIe
98%16.7
NVIDIA A10G
38%6.4
NVIDIA L4
29%5
NVIDIA T4
17%2.9

← All AI & Machine Learning GPU rankings

What this card can build

Whole-job timings, composed from our measured per-model results on this card.

WorkflowTimeEnergyBasis
50-image depth pass10 s0.12 Whmeasured
30-minute podcast pass59 s2.13 Whall 3 stages measured
24-frame storyboard3 min2.08 Whestimate, 1 of 2 stages measured
500-image masking run3.5 min18.04 Whmeasured
60-second AI short film4 min6.39 Whestimate, 2 of 3 stages measured
60-second AI short film, narrated4.2 min6.4 Whestimate, 3 of 4 stages measured
6-panel comic page4.8 min1.04 Whestimate, 1 of 3 stages measured
Character sheet, 12 poses6.3 minn/aestimate, 0 of 2 stages measured
Full codebase review11.7 min35.27 Whmeasured
200-product catalogue cutout12.2 min22.9 Whall 2 stages measured
10 short social clips13.7 min0.74 Whestimate, 1 of 3 stages measured
40-product photo shoot22.2 min13.47 Whestimate, 1 of 2 stages measured
40-product shoot, start to finish24.8 min18.05 Whestimate, 3 of 4 stages measured
100-photo restoration batch50.2 minn/aestimate, 0 of 1 stage measured
100-photo restore and enlarge56 min11.25 Whestimate, 1 of 2 stages measured

Can't run: 20 long-form articles (needs Llama-3.3-70B).

Rent or buy?

This card is $12,000 to buy. The cheapest listed rate on RunPod is $1.000/hour, but that is the floor: we budget $1.200/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 10,000 GPU-hours. Below it you are paying for idle silicon.

How you would use itGPU-hours a yearRental cost a yearTime to break even
2 hours a day, hobby730$87613.7 years
8 hours a day, working on it2,920$3,5043.4 years
24/7, always-on agent8,760$10,5121.1 years

At hobby usage this card is very unlikely to pay for itself before it is superseded. Rent it. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.

Rental price

$1.000/hr+0.0% since 2026-08-14low $1.000 · high $1.000

Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.