32GB · AI Score 11.0/100 · first-party measured on 12 AI workloads
11 AI Score ✓ Measured
Every number on this page is first-party: NVIDIA RTX 5000 Ada Generation was run on our pinned 12-workload AI suite on 2026-07-12, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) NVIDIA RTX 5000 Ada Generation delivers about 103.18 tokens/sec. Stepping up to Qwen3 32B it holds roughly 25.88 tok/s. Llama 3.3 70B does not fit. It needs roughly 42GB and this card has 32GB. For image generation, SDXL runs at 6.37 it/s, and FLUX.1-dev at 0.6 it/s. 2 of the 12 workloads won't fit on 32GB at the tested precision, Llama 3.3 70B, Qwen-Image-Edit. We publish those as hard gates rather than quietly dropping to a smaller quant.
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Qwen3 4B | 170.72 tok/s | 2.8 GB peak157 W65°C1.09 tok/WQ4_K_M | ✓ Measured |
| Llama 3.1 8B | 103.18 tok/s | 5 GB peak201 W70°C0.51 tok/WQ4_K_M | ✓ Measured |
| Qwen2.5-Coder 14B | 56.66 tok/s | 8.7 GB peak202 W77°C0.28 tok/WQ4_K_M | ✓ Measured |
| Qwen3 32B | 25.88 tok/s | 18.8 GB peak197 W83°C0.13 tok/WQ4_K_M | ✓ Measured |
| Llama 3.3 70B | ✕ Won't fit needs ~46 GB | VRAM-gated at this precision | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Stable Diffusion XL | 12.74 images/min | 14.7 GB peak236 W84°C4.7 s/img | ✓ Measured |
| FLUX.1 dev | 1.29 images/min | 23.6 GB peak149 W86°C46.9 s/img | ✓ Measured CPU offload |
| Z-Image Turbo | 5.93 images/min | 25.7 GB peak229 W86°C10.2 s/img | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| FLUX.1 Kontext dev | 0.79 images/min | 24.8 GB peak178 W89°C76.5 s/img | ✓ Measured CPU offload |
| Qwen-Image-Edit | ✕ Won't fit needs ~42 GB | VRAM-gated at this precision | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| LTX-Video (distilled) | 3.23 frames/s | 9.3 GB peak164 W85°C30 s/clip | ✓ Measured CPU offload |
| Wan 2.2 5B (720p) | 0.32 frames/s | 18.5 GB peak211 W90°C153.1 s/clip | ✓ Measured CPU offload |
| Architecture | Ada Lovelace |
| CUDA cores | 12,800 |
| VRAM | 32GB GDDR6 ECC |
| Memory bus | 256-bit |
| Memory bandwidth | 576 GB/s |
| Boost clock | 2,550 MHz |
| TDP | 250 W |
| Process | 4nm |
| Interface | PCIe 4.0 x16 |
| Release date | 2023-08-09 |
| Launch MSRP | $4,000 |
NVIDIA RTX 5000 Ada Generation scores 11.0/100, #27 of 102. It ran 10 of 12; 2 exceeded its 32GB. Every figure here is our own measurement.
100% = this card, AI & Machine Learning headline metric (AI Score). #4 of 61 desktop cards in this vertical.
| GPU | Relative | % | AI Score |
|---|---|---|---|
| NVIDIA RTX 6000 Ada Generation | 216% | 23.8 | |
| NVIDIA GeForce RTX 5090 | 203% | 22.3 | |
| NVIDIA RTX 5880 Ada Generation | 134% | 14.7 | |
| NVIDIA RTX 5000 Ada Generation | 100% | 11 | |
| NVIDIA GeForce RTX 4090 | 96% | 10.6 | |
| NVIDIA GeForce RTX 3090 Ti | 77% | 8.5 | |
| NVIDIA Titan RTX | 75% | 8.2 | |
| NVIDIA GeForce RTX 3090 | 70% | 7.7 |
← All AI & Machine Learning GPU rankings
| Transistors | 76,300 million |
| Die size | 608.4 mm² |
| Process node | 4 nm |
| Fabricated by | TSMC |
| Transistor density | 125.4 million per mm² |
Denser than 97% of the 746 cards we have silicon data for. Density is the clearest measure of what a process node bought: a card that gained it without growing the die got its speed from the fab rather than the architecture.
Silicon figures from Wikipedia (CC BY-SA 4.0). Benchmarks on this page are our own. Compare every chip.
Whole-job timings, composed from our measured per-model results on this card.
| Workflow | Time | Energy | Basis |
|---|---|---|---|
| 24-frame storyboard | 5.1 min | 18.41 Wh | all 2 stages measured |
| 60-second AI short film | 9.5 min | 27.06 Wh | all 3 stages measured |
| 6-panel comic page | 12.8 min | 35.55 Wh | all 3 stages measured |
| Character sheet, 12 poses | 16 min | 46.88 Wh | all 2 stages measured |
| Full codebase review | 17.6 min | 59.48 Wh | measured |
| Short social clips | 27.9 min | 97.41 Wh | all 3 stages measured |
| Product photo shoot | 53.7 min | 162.13 Wh | all 2 stages measured |
| Photo restoration batch | 2 h 6 min | 374.53 Wh | measured |
Can't run: Long-form article batch (needs Llama 3.3 70B).
This card is $4,000 to buy. The cheapest listed rate on Vast.ai is $0.388/hour, but that is the floor: we budget $0.466/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 8,591 GPU-hours. Below it you are paying for idle silicon.
| How you would use it | GPU-hours a year | Rental cost a year | Time to break even |
|---|---|---|---|
| 2 hours a day, hobby | 730 | $340 | 11.8 years |
| 8 hours a day, working on it | 2,920 | $1,360 | 2.9 years |
| 24/7, always-on agent | 8,760 | $4,079 | 11.8 months |
At hobby usage this card is very unlikely to pay for itself before it is superseded. Rent it. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.
Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.