48GB · AI Score 17.7/100 · first-party measured on 12 AI workloads
17.7 AI Score Includes estimates
Every number on this page is first-party: NVIDIA A40 was run on our pinned 12-workload AI suite on 2026-07-11, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) NVIDIA A40 delivers about 106.18 tokens/sec. Stepping up to Qwen3 32B it holds roughly 26.89 tok/s. Llama 3.3 70B does not fit. It needs roughly 42GB and this card has 48GB. For image generation, SDXL runs at 4.6 it/s, and FLUX.1-dev at 0.91 it/s. 1 of the 12 workloads won't fit on 48GB at the tested precision, Llama 3.3 70B. We publish those as hard gates rather than quietly dropping to a smaller quant. NVIDIA A40 isn't a retail purchase for most people. It's rented by the hour. You can run this exact card on RunPod.
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Qwen3 4B | 160.79 tok/s | 3 GB peak165 W49°C0.97 tok/WQ4_K_M | ✓ Measured |
| Llama 3.1 8B | 106.18 tok/s | 4.9 GB peak182 W53°C0.59 tok/WQ4_K_M | ✓ Measured |
| Qwen2.5-Coder 14B | 57.85 tok/s | 8.7 GB peak188 W58°C0.31 tok/WQ4_K_M | ✓ Measured |
| Qwen3 32B | 26.89 tok/s | 18.7 GB peak170 W59°C0.16 tok/WQ4_K_M | ✓ Measured |
| Llama 3.3 70B | n/a tok/s | estimated | Est. |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Stable Diffusion XL | 9.2 images/min | 15.8 GB peak295 W57°C6.5 s/img | ✓ Measured |
| Z-Image Turbo | 3.23 images/min | 44.3 GB peak298 W66°C18.5 s/img | ✓ Measured |
| FLUX.1 dev | 1.95 images/min | 44.4 GB peak297 W68°C30.7 s/img | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Wan 2.2 5B (720p) | 0.31 frames/s | 44.4 GB peak294 W69°C159.7 s/clip | ✓ Measured |
| Architecture | Ampere |
| CUDA cores | 10,752 |
| VRAM | 48GB GDDR6 |
| Memory bus | 384-bit |
| Memory bandwidth | 696 GB/s |
| Boost clock | 1,740 MHz |
| TDP | 300 W |
| Process | Samsung 8nm |
| Interface | PCIe 4.0 x16 |
| Release date | 2020-10-05 |
| Launch MSRP | $5,000 |
NVIDIA A40 scores 17.7/100, #22 of 102. It ran 8 of 12; 1 exceeded its 48GB. Every figure here is our own measurement.
100% = this card, AI & Machine Learning headline metric (AI Score). #16 of 21 desktop cards in this vertical.
| GPU | Relative | % | AI Score |
|---|---|---|---|
| NVIDIA A800 80GB | 187% | 33.1 | |
| NVIDIA A100 80GB PCIe | 179% | 31.7 | |
| NVIDIA L40S | 156% | 27.7 | |
| NVIDIA L40 | 110% | 19.5 | |
| NVIDIA A40 | 100% | 17.7 | |
| NVIDIA A100 40GB SXM4 | 96% | 17 | |
| NVIDIA A100 40GB PCIe | 94% | 16.7 | |
| NVIDIA A10G | 36% | 6.4 | |
| NVIDIA L4 | 28% | 5 |
← All AI & Machine Learning GPU rankings
Whole-job timings, composed from our measured per-model results on this card.
| Workflow | Time | Energy | Basis |
|---|---|---|---|
| 24-frame storyboard | 10.2 min | 39.46 Wh | all 2 stages measured |
| Full codebase review | 17.3 min | 54.16 Wh | measured |
| 10 short social clips | 33.4 min | 145.56 Wh | all 3 stages measured |
This card is $5,000 to buy. The cheapest listed rate on RunPod is $0.350/hour, but that is the floor: we budget $0.420/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 11,905 GPU-hours. Below it you are paying for idle silicon.
| How you would use it | GPU-hours a year | Rental cost a year | Time to break even |
|---|---|---|---|
| 2 hours a day, hobby | 730 | $307 | 16.3 years |
| 8 hours a day, working on it | 2,920 | $1,226 | 4.1 years |
| 24/7, always-on agent | 8,760 | $3,679 | 1.4 years |
At hobby usage this card is very unlikely to pay for itself before it is superseded. Rent it. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.
Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.