40GB · AI Score 17.0/100 · anchored estimate vs 51 measured cards
17 AI Score Includes estimates
We have not run NVIDIA A100 40GB SXM4 on our bench. These figures are anchored estimates, interpolated per workload against the 51 GPUs we did measure (confidence: high (sibling silicon)). On Llama 3.1 8B (Q4_K_M) NVIDIA A100 40GB SXM4 should deliver about 124 tokens/sec. Stepping up to Qwen3 32B it should hold roughly 34.7 tok/s. Llama 3.3 70B does not fit. It needs roughly 42GB and this card has 40GB. For image generation, SDXL should run near 8.18 it/s, and FLUX.1-dev at 1.99 it/s. 2 of the 12 workloads won't fit on 40GB at the tested precision, Llama 3.3 70B, Qwen-Image-Edit. We publish those as hard gates rather than quietly dropping to a smaller quant.
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| gemma-3-270m | 574.24 tok/s | 74 W36°CQ4_K_M | ✓ Measured |
| LFM2.5-1.2B | 572.69 tok/s | 94 W42°CQ4_K_M | ✓ Measured |
| SmolLM2-135M | 544.06 tok/s | 69 W38°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder-0.5B | 542.36 tok/s | 75 W37°CQ4_K_M | ✓ Measured |
| Llama 3.2 1B | 532.08 tok/s | 113 W41°CQ4_K_M | ✓ Measured |
| Qwen2.5-0.5B | 517.09 tok/s | 69 W36°CQ4_K_M | ✓ Measured |
| Qwen1.5-0.5B | 497.73 tok/s | 71 W36°CQ4_K_M | ✓ Measured |
| Qwen3 0.6B | 417.88 tok/s | 70 W36°CQ4_K_M | ✓ Measured |
| Qwen3-0.6B | 408.39 tok/s | 76 W38°CQ4_K_M | ✓ Measured |
| LFM2.5-8B-A1B | 351.47 tok/s | 93 W42°CQ4_K_M | ✓ Measured |
| Qwen3-1.7B | 341.85 tok/s | 103 W41°CQ4_K_M | ✓ Measured |
| Qwen3 1.7B | 340.76 tok/s | 112 W40°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder-1.5B | 322.96 tok/s | 95 W39°CQ4_K_M | ✓ Measured |
| Qwen2-1.5B | 322.09 tok/s | 93 W41°CQ4_K_M | ✓ Measured |
| DeepSeek-R1 Distill 1.5B | 320.78 tok/s | 93 W41°CQ4_K_M | ✓ Measured |
| Qwen2.5-1.5B | 315.59 tok/s | 100 W40°CQ4_K_M | ✓ Measured |
| gemma-3-1b | 310.69 tok/s | 88 W38°CQ4_K_M | ✓ Measured |
| Llama 3.2 3B | 261.24 tok/s | 118 W44°CQ4_K_M | ✓ Measured |
| Hermes-3-Llama-3.2-3B | 261 tok/s | 126 W44°CQ4_K_M | ✓ Measured |
| Llama-3.2-3B-Instruct-uncensored | 260.83 tok/s | 115 W43°CQ4_K_M | ✓ Measured |
| SmolLM3-3B | 247.66 tok/s | 131 W43°CQ4_K_M | ✓ Measured |
| SmolLM3 3B | 246.25 tok/s | 129 W43°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder-3B | 241.51 tok/s | 129 W43°CQ4_K_M | ✓ Measured |
| Qwen2.5-3B | 240.48 tok/s | 127 W43°CQ4_K_M | ✓ Measured |
| Phi-4-mini | 237.38 tok/s | 125 W46°CQ4_K_M | ✓ Measured |
| Phi-4 Mini 3.8B | 236.71 tok/s | 139 W46°CQ4_K_M | ✓ Measured |
| gemma-2-2b | 234.68 tok/s | 109 W43°CQ4_K_M | ✓ Measured |
| gemma-2-2b-it-abliterated | 232.27 tok/s | 121 W43°CQ4_K_M | ✓ Measured |
| AI21-Jamba-Reasoning-3B | 231.15 tok/s | 110 W43°CQ4_K_M | ✓ Measured |
| phi-2 | 228.1 tok/s | 131 W43°CQ4_K_M | ✓ Measured |
| Phi-3.5-mini | 222.53 tok/s | 132 W44°CQ4_K_M | ✓ Measured |
| DeepSeek-Coder-V2-Lite | 205.42 tok/s | 120 W43°CQ4_K_M | ✓ Measured |
| gpt-oss-20b | 205.1 tok/s | 121 W45°CQ4_K_M | ✓ Measured |
| Qwen3-4B-Instruct-2507 | 193.39 tok/s | 140 W45°CQ4_K_M | ✓ Measured |
| Qwen3-4B | 193.14 tok/s | 137 W44°CQ4_K_M | ✓ Measured |
| Qwen3-4B-Instruct-2507 | 192.89 tok/s | 133 W45°CQ4_K_M | ✓ Measured |
| Nemotron-3-Nano-30B-A3B | 192.44 tok/s | 103 W43°CQ4_K_M | ✓ Measured |
| Qwen3-4B-Thinking-2507 | 189.89 tok/s | 134 W44°CQ4_K_M | ✓ Measured |
| Gemma 3 4B | 174.14 tok/s | 138 W43°CQ4_K_M | ✓ Measured |
| Qwen3-Coder 30B A3B | 173.47 tok/s | 92 W42°CQ4_K_M | ✓ Measured |
| Llama-2-7B | 172.53 tok/s | 195 W49°CQ4_K_M | ✓ Measured |
| Qwen3 30B A3B | 169.1 tok/s | 92 W42°CQ4_K_M | ✓ Measured |
| Qwen3-30B-A3B | 167.1 tok/s | 106 W42°CQ4_K_M | ✓ Measured |
| Mistral-7B-Instruct-v0.3 | 167.06 tok/s | 167 W47°CQ4_K_M | ✓ Measured |
| Mistral-7B-Instruct-v0.2 | 166.8 tok/s | 165 W47°CQ4_K_M | ✓ Measured |
| Mistral-7B-Instruct-v0.1 | 166.39 tok/s | 174 W48°CQ4_K_M | ✓ Measured |
| Mistral 7B v0.3 | 165.9 tok/s | 182 W48°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder 7B | 159.9 tok/s | 176 W47°CQ4_K_M | ✓ Measured |
| Qwen2.5-7B | 159.81 tok/s | 172 W48°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder-7B-Instruct-abliterated | 159.76 tok/s | 151 W47°CQ4_K_M | ✓ Measured |
| DeepSeek-R1 Distill 7B | 159.09 tok/s | 164 W47°CQ4_K_M | ✓ Measured |
| Llama-3.1-8B | 157.26 tok/s | 170 W47°CQ4_K_M | ✓ Measured |
| DeepSeek-R1 Distill Llama 8B | 157.16 tok/s | 166 W48°CQ4_K_M | ✓ Measured |
| Meta-Llama-3.1-8B | 156.68 tok/s | 170 W48°CQ4_K_M | ✓ Measured |
| DarkIdol-Llama-3.1-8B-Instruct-1.2-Uncensored | 156.62 tok/s | 177 W47°CQ4_K_M | ✓ Measured |
| Dolphin 3.0 Llama 3.1 8B | 156.55 tok/s | 138 W49°CQ4_K_M | ✓ Measured |
| L3-8B-Stheno-v3.2 | 156.55 tok/s | 163 W49°CQ4_K_M | ✓ Measured |
| Dolphin X1 8B | 156.3 tok/s | 153 W49°CQ4_K_M | ✓ Measured |
| dolphin-2.9-llama3-8b | 153.34 tok/s | 140 W48°CQ4_K_M | ✓ Measured |
| Dolphin X1 Trinity Nano 6B | 153.27 tok/s | 91 W39°CQ4_K_M | ✓ Measured |
| Josiefied-Qwen3-8B-abliterated-v1 | 146.8 tok/s | 168 W47°CQ4_K_M | ✓ Measured |
| Qwen3 8B | 146.21 tok/s | 164 W46°CQ4_K_M | ✓ Measured |
| Qwen3-8B | 145.96 tok/s | 165 W47°CQ4_K_M | ✓ Measured |
| DeepSeek-R1-0528-Qwen3-8B | 145.84 tok/s | 142 W47°CQ4_K_M | ✓ Measured |
| Qwen3 30B A3B (Q3_K_M) | 142.78 tok/s | 115 W43°CQ3_K_M | ✓ Measured |
| KAT-Coder-V2.5-Dev | 137.31 tok/s | 107 W42°CQ4_K_M | ✓ Measured |
| Ornith-1.0-35B | 131.77 tok/s | 104 W41°CQ4_K_M | ✓ Measured |
| Qwen-AgentWorld-35B-A3B | 131.16 tok/s | 97 W43°CQ4_K_M | ✓ Measured |
| Ornith-1.0-9B | 126.79 tok/s | 170 W48°CQ4_K_M | ✓ Measured |
| GLM-4.7-Flash | 119.1 tok/s | 98 W41°CQ4_K_M | ✓ Measured |
| GLM-4.7-Flash-REAP-23B-A3B | 109.1 tok/s | 103 W42°CQ4_K_M | ✓ Measured |
| Mistral-Nemo-Instruct-2407 | 106.33 tok/s | 171 W49°CQ4_K_M | ✓ Measured |
| NemoMix-Unleashed-12B | 105.05 tok/s | 161 W49°CQ4_K_M | ✓ Measured |
| gemma-2-9b | 101.97 tok/s | 175 W47°CQ4_K_M | ✓ Measured |
| Phi-4 14B | 98.65 tok/s | 200 W52°CQ4_K_M | ✓ Measured |
| Qwen3-14B | 90.07 tok/s | 188 W50°CQ4_K_M | ✓ Measured |
| Qwen3 14B | 90.01 tok/s | 206 W50°CQ4_K_M | ✓ Measured |
| Hermes-4-14B | 89.97 tok/s | 166 W50°CQ4_K_M | ✓ Measured |
| Gemma 3 12B | 88.97 tok/s | 175 W48°CQ4_K_M | ✓ Measured |
| Gemma 4 12B | 87.14 tok/s | 162 W47°CQ4_K_M | ✓ Measured |
| EVA-Qwen2.5-14B-v0.2 | 86.48 tok/s | 187 W50°CQ4_K_M | ✓ Measured |
| DeepSeek-R1 Distill 14B | 86.39 tok/s | 167 W50°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder-14B-Instruct-abliterated | 86.34 tok/s | 183 W49°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder-14B | 86.2 tok/s | 182 W49°CQ4_K_M | ✓ Measured |
| Qwen2.5-14B | 86.07 tok/s | 193 W49°CQ4_K_M | ✓ Measured |
| Uncensored | 85.82 tok/s | 188 W50°CQ4_K_M | ✓ Measured |
| Phi-4 14B (Q3_K_M) | 80.71 tok/s | 182 W50°CQ3_K_M | ✓ Measured |
| StarCoder2 15B | 80.25 tok/s | 217 W52°CQ4_K_M | ✓ Measured |
| Gemma 3 12B (Q3_K_M) | 69.7 tok/s | 183 W48°CQ3_K_M | ✓ Measured |
| DeepSeek-R1 Distill 14B (Q3_K_M) | 65.33 tok/s | 187 W49°CQ3_K_M | ✓ Measured |
| Codestral 22B | 63.74 tok/s | 203 W52°CQ4_K_M | ✓ Measured |
| Dolphin-Mistral-24B-Venice-Edition | 62.56 tok/s | 212 W52°CQ4_K_M | ✓ Measured |
| Dolphin Mistral 24B Venice | 62.46 tok/s | 214 W53°CQ4_K_M | ✓ Measured |
| Mistral Small 24B | 62.42 tok/s | 204 W53°CQ4_K_M | ✓ Measured |
| Cydonia-24B-v4.3 | 62.41 tok/s | 195 W53°CQ4_K_M | ✓ Measured |
| Devstral Small 24B | 62.4 tok/s | 196 W52°CQ4_K_M | ✓ Measured |
| Dolphin 3.0 R1 Mistral 24B | 62.4 tok/s | 204 W52°CQ4_K_M | ✓ Measured |
| Codestral 22B (Q3_K_M) | 48.37 tok/s | 210 W52°CQ3_K_M | ✓ Measured |
| Gemma 3 27B | 47.47 tok/s | 200 W52°CQ4_K_M | ✓ Measured |
| Mistral Small 24B (Q3_K_M) | 45.34 tok/s | 200 W52°CQ3_K_M | ✓ Measured |
| Qwen3-32B | 43.8 tok/s | 235 W54°CQ4_K_M | ✓ Measured |
| Olmo-3.1-32B-Think | 43.67 tok/s | 212 W53°CQ4_K_M | ✓ Measured |
| Qwen2.5-32B | 43.23 tok/s | 206 W53°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder 32B | 43.14 tok/s | 208 W53°CQ4_K_M | ✓ Measured |
| DeepSeek-R1 Distill 32B | 43.13 tok/s | 212 W54°CQ4_K_M | ✓ Measured |
| DeepSeek-R1-Distill-Qwen-32B-abliterated | 43.07 tok/s | 192 W53°CQ4_K_M | ✓ Measured |
| QwQ 32B | 43.02 tok/s | 220 W53°CQ4_K_M | ✓ Measured |
| Dolphin 2.9.1 Yi 1.5 34B | 42.97 tok/s | 233 W54°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder 32B (Q3_K_M) | 31.04 tok/s | 210 W53°CQ3_K_M | ✓ Measured |
| Llama-3.3-70B | ✕ Won't fit needs ~54 GB | VRAM-gated at this precision | ✓ Measured |
| Qwen3-Coder-Next-abliterated | ✕ Won't fit needs ~61 GB | VRAM-gated at this precision | ✓ Measured |
| Laguna-XS-2.1 | ✕ Won't fit needs ~26 GB | VRAM-gated at this precision | ✓ Measured |
| Nanbeige4.2-3B | ✕ Won't fit needs ~4 GB | VRAM-gated at this precision | ✓ Measured |
| Qwen3-Coder-Next | ✕ Won't fit needs ~61 GB | VRAM-gated at this precision | ✓ Measured |
| Qwen3-Next-80B-A3B-Thinking | ✕ Won't fit needs ~61 GB | VRAM-gated at this precision | ✓ Measured |
| Qwen2.5-72B | ✕ Won't fit needs ~60 GB | VRAM-gated at this precision | ✓ Measured |
| Qwen3-Coder-Next | ✕ Won't fit needs ~61 GB | VRAM-gated at this precision | ✓ Measured |
| Qwen3-Next-80B-A3B-Thinking | ✕ Won't fit needs ~61 GB | VRAM-gated at this precision | ✓ Measured |
| Qwen3-Next-80B-A3B | ✕ Won't fit needs ~61 GB | VRAM-gated at this precision | ✓ Measured |
| DeepSeek-R1-Distill-Llama-70B | ✕ Won't fit needs ~54 GB | VRAM-gated at this precision | ✓ Measured |
| Hermes-4-70B | ✕ Won't fit needs ~54 GB | VRAM-gated at this precision | ✓ Measured |
| Llama-3.3-70B-Instruct-abliterated | ✕ Won't fit needs ~54 GB | VRAM-gated at this precision | ✓ Measured |
| Meta-Llama-3.1-70B | ✕ Won't fit needs ~54 GB | VRAM-gated at this precision | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| FLUX.1 Schnell | 28.44 images/min | 384 W56°C | ✓ Measured |
| Z-Image Turbo (1024px) | 18.7 images/min | 382 W61°C | ✓ Measured |
| Stable Diffusion XL | 18.58 images/min | 10.3 GB peak375 W61°C3.2 s/img | ✓ Measured |
| Z-Image Turbo | 9.98 images/min | estimated | Est. |
| FLUX.1 dev | 4.26 images/min | estimated | Est. |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| TinyLlama 1.1B LoRA | 7471 train tok/s | 189 W57°C | ✓ Measured |
| SmolLM2 1.7B LoRA | 6938.8 train tok/s | 254 W60°C | ✓ Measured |
| Qwen2.5 1.5B LoRA | 6349.7 train tok/s | 210 W58°C | ✓ Measured |
| Qwen2.5 7B LoRA | 3670.8 train tok/s | 381 W67°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Qwen2.5 1.5B served | 4111.1 serve tok/s | 178 W45°C | ✓ Measured |
| SmolLM2 1.7B served | 3865.2 serve tok/s | 187 W45°C | ✓ Measured |
| TinyLlama 1.1B served | 3221.3 serve tok/s | 153 W40°C | ✓ Measured |
| Qwen2.5 7B served | 1769.6 serve tok/s | 225 W49°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| FLUX.1 Kontext dev | 1.99 images/min | estimated | Est. |
| Qwen-Image-Edit | ✕ Won't fit | VRAM-gated at this precision | Est. |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| LTX-Video (distilled) | 8.95 frames/s | estimated | Est. |
| Wan 2.2 5B (720p) | 0.66 frames/s | estimated | Est. |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Depth Anything V2 Small | 613.42 images/min | 66 W35°C | ✓ Measured |
| Depth Anything V2 Large | 536.74 images/min | 74 W35°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| SAM ViT-Base | 696.77 images/min | 64 W37°C | ✓ Measured |
| SAM ViT-Huge | 146.81 images/min | 318 W52°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| BiRefNet | 611.56 images/min | 76 W35°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Swin2SR 4x Upscaler | 17.25 images/min | 116 W44°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Kokoro TTS 82M | 142.03 x realtime | 60 W32°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| MusicGen Small | 1.08 x realtime | 94 W34°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Whisper large-v3 | 102.47 x realtime | 190 W36°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Qwen2.5 1.5B + 0.5B draft | 0.69 x vs solo | ✓ Measured |
| Architecture | Ampere |
| CUDA cores | 6,912 |
| VRAM | 40GB HBM2 |
| Memory bus | 5120-bit |
| Memory bandwidth | 1555 GB/s |
| Boost clock | 1,410 MHz |
| TDP | 400 W |
| Process | TSMC 7nm |
| Interface | SXM4 |
| Release date | 2020-05-14 |
| Launch MSRP | $12,000 |
NVIDIA A100 40GB SXM4 scores 17.0/100, #23 of 102. It ran 10 of 12; 2 exceeded its 40GB. Figures are anchored estimates, not measurements, we flag that on every row.
100% = this card, AI & Machine Learning headline metric (AI Score). #17 of 21 desktop cards in this vertical.
| GPU | Relative | % | AI Score |
|---|---|---|---|
| NVIDIA A100 80GB PCIe | 186% | 31.7 | |
| NVIDIA L40S | 163% | 27.7 | |
| NVIDIA L40 | 115% | 19.5 | |
| NVIDIA A40 | 104% | 17.7 | |
| NVIDIA A100 40GB SXM4 | 100% | 17 | |
| NVIDIA A100 40GB PCIe | 98% | 16.7 | |
| NVIDIA A10G | 38% | 6.4 | |
| NVIDIA L4 | 29% | 5 | |
| NVIDIA T4 | 17% | 2.9 |
← All AI & Machine Learning GPU rankings
Whole-job timings, composed from our measured per-model results on this card.
| Workflow | Time | Energy | Basis |
|---|---|---|---|
| 50-image depth pass | 10 s | 0.12 Wh | measured |
| 30-minute podcast pass | 59 s | 2.13 Wh | all 3 stages measured |
| 24-frame storyboard | 3 min | 2.08 Wh | estimate, 1 of 2 stages measured |
| 500-image masking run | 3.5 min | 18.04 Wh | measured |
| 60-second AI short film | 4 min | 6.39 Wh | estimate, 2 of 3 stages measured |
| 60-second AI short film, narrated | 4.2 min | 6.4 Wh | estimate, 3 of 4 stages measured |
| 6-panel comic page | 4.8 min | 1.04 Wh | estimate, 1 of 3 stages measured |
| Character sheet, 12 poses | 6.3 min | n/a | estimate, 0 of 2 stages measured |
| Full codebase review | 11.7 min | 35.27 Wh | measured |
| 200-product catalogue cutout | 12.2 min | 22.9 Wh | all 2 stages measured |
| 10 short social clips | 13.7 min | 0.74 Wh | estimate, 1 of 3 stages measured |
| 40-product photo shoot | 22.2 min | 13.47 Wh | estimate, 1 of 2 stages measured |
| 40-product shoot, start to finish | 24.8 min | 18.05 Wh | estimate, 3 of 4 stages measured |
| 100-photo restoration batch | 50.2 min | n/a | estimate, 0 of 1 stage measured |
| 100-photo restore and enlarge | 56 min | 11.25 Wh | estimate, 1 of 2 stages measured |
Can't run: 20 long-form articles (needs Llama-3.3-70B).
This card is $12,000 to buy. The cheapest listed rate on RunPod is $1.000/hour, but that is the floor: we budget $1.200/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 10,000 GPU-hours. Below it you are paying for idle silicon.
| How you would use it | GPU-hours a year | Rental cost a year | Time to break even |
|---|---|---|---|
| 2 hours a day, hobby | 730 | $876 | 13.7 years |
| 8 hours a day, working on it | 2,920 | $3,504 | 3.4 years |
| 24/7, always-on agent | 8,760 | $10,512 | 1.1 years |
At hobby usage this card is very unlikely to pay for itself before it is superseded. Rent it. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.
Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.