24GB · AI Score 7.4/100 · first-party measured on 12 AI workloads
7.4 AI Score ✓ Measured
Every number on this page is first-party: NVIDIA RTX A5000 was run on our pinned 12-workload AI suite on 2026-07-12, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) NVIDIA RTX A5000 delivers about 119.86 tokens/sec. Stepping up to Qwen3 32B it holds roughly 30.49 tok/s. Llama 3.3 70B does not fit. It needs roughly 42GB and this card has 24GB. For image generation, SDXL runs at 3.73 it/s, while FLUX.1-dev won't fit at BF16 (needs ~26GB). 4 of the 12 workloads won't fit on 24GB at the tested precision, Llama 3.3 70B, FLUX.1-dev, FLUX.1 Kontext, Qwen-Image-Edit. We publish those as hard gates rather than quietly dropping to a smaller quant. NVIDIA RTX A5000 isn't a retail purchase for most people. It's rented by the hour. You can run this exact card on RunPod.
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Qwen3 4B | 171.68 tok/s | 2.7 GB peak136 W53°C1.26 tok/WQ4_K_M | ✓ Measured |
| Llama 3.1 8B | 119.86 tok/s | 4.9 GB peak154 W58°C0.78 tok/WQ4_K_M | ✓ Measured |
| Qwen2.5-Coder 14B | 64.89 tok/s | 8.6 GB peak180 W65°C0.36 tok/WQ4_K_M | ✓ Measured |
| Qwen3 32B | 30.49 tok/s | 18.7 GB peak161 W73°C0.19 tok/WQ4_K_M | ✓ Measured |
| Llama 3.3 70B | ✕ Won't fit needs ~46 GB | VRAM-gated at this precision | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Stable Diffusion XL | 7.46 images/min | 15.7 GB peak227 W80°C8 s/img | ✓ Measured |
| Z-Image Turbo | 4.05 images/min | 23.1 GB peak228 W86°C14.7 s/img | ✓ Measured |
| FLUX.1 dev | ✕ Won't fit needs ~26 GB | VRAM-gated at this precision | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| FLUX.1 Kontext dev | ✕ Won't fit needs ~26 GB | VRAM-gated at this precision | ✓ Measured |
| Qwen-Image-Edit | ✕ Won't fit needs ~42 GB | VRAM-gated at this precision | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| LTX-Video (distilled) | 2.7 frames/s | 9.2 GB peak182 W85°C35.9 s/clip | ✓ Measured |
| Wan 2.2 5B (720p) | 0.24 frames/s | 16.6 GB peak211 W87°C201.2 s/clip | ✓ Measured |
| Architecture | Ampere (GA102 derivative) |
| CUDA cores | 8,192 |
| VRAM | 24GB GDDR6 |
| Memory bus | 384-bit |
| Memory bandwidth | 768 GB/s |
| Boost clock | 1,695 MHz |
| TDP | 230 W |
| Process | 8nm |
| Interface | PCIe 4.0 x16 |
| Release date | 2021-04-12 |
| Launch MSRP | $2,250 |
NVIDIA RTX A5000 scores 7.4/100, #37 of 102. It ran 8 of 12; 4 exceeded its 24GB. Every figure here is our own measurement.
100% = this card, AI & Machine Learning headline metric (AI Score). #11 of 20 desktop cards in this vertical.
| GPU | Relative | % | AI Score |
|---|---|---|---|
| NVIDIA Quadro RTX 8000 | 111% | 8.2 | |
| NVIDIA RTX A5500 | 103% | 7.6 | |
| NVIDIA RTX PRO 4000 Blackwell | 103% | 7.6 | |
| NVIDIA Quadro RTX 6000 (Turing) | 100% | 7.4 | |
| NVIDIA RTX A5000 | 100% | 7.4 | |
| AMD Radeon Pro W7800 | 84% | 6.2 | |
| NVIDIA RTX 4500 Ada Generation | 80% | 5.9 | |
| AMD Radeon Pro W6800 | 74% | 5.5 | |
| NVIDIA RTX A4500 | 69% | 5.1 |
← All AI & Machine Learning GPU rankings
Whole-job timings, composed from our measured per-model results on this card.
| Workflow | Time | Energy | Basis |
|---|---|---|---|
| 24-frame storyboard | 7 min | 24.57 Wh | all 2 stages measured |
| 60-second AI short film | 11.8 min | 36.17 Wh | all 3 stages measured |
| Full codebase review | 15.4 min | 46.13 Wh | measured |
| 10 short social clips | 37.2 min | 129.95 Wh | all 3 stages measured |
Can't run: 40-product photo shoot (needs FLUX.1 Kontext dev), 6-panel comic page (needs FLUX.1 dev), 20 long-form articles (needs Llama 3.3 70B), Character sheet, 12 poses (needs FLUX.1 dev), 100-photo restoration batch (needs FLUX.1 Kontext dev), 100-photo restore and enlarge (needs FLUX.1 Kontext dev).
This card is $2,250 to buy. The cheapest listed rate on RunPod is $0.160/hour, but that is the floor: we budget $0.192/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 11,719 GPU-hours. Below it you are paying for idle silicon.
| How you would use it | GPU-hours a year | Rental cost a year | Time to break even |
|---|---|---|---|
| 2 hours a day, hobby | 730 | $140 | 16.1 years |
| 8 hours a day, working on it | 2,920 | $561 | 4.0 years |
| 24/7, always-on agent | 8,760 | $1,682 | 1.3 years |
At hobby usage this card is very unlikely to pay for itself before it is superseded. Rent it. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.
Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.