32GB · AI Score 22.3/100 · first-party measured on 12 AI workloads
22.3 AI Score ✓ Measured
Every number on this page is first-party: NVIDIA GeForce RTX 5090 was run on our pinned 12-workload AI suite on 2026-07-11, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) NVIDIA GeForce RTX 5090 delivers about 268.14 tokens/sec. Stepping up to Qwen3 32B it holds roughly 71.15 tok/s. Llama 3.3 70B does not fit. It needs roughly 42GB and this card has 32GB. For image generation, SDXL runs at 10.54 it/s, and FLUX.1-dev at 0.95 it/s. 2 of the 12 workloads won't fit on 32GB at the tested precision, Llama 3.3 70B, Qwen-Image-Edit. We publish those as hard gates rather than quietly dropping to a smaller quant. NVIDIA GeForce RTX 5090 isn't a retail purchase for most people. It's rented by the hour. You can run this exact card on RunPod.
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Qwen3 4B | 375.6 tok/s | 3 GB peak87 W48°C4.32 tok/WQ4_K_M | ✓ Measured |
| Llama 3.1 8B | 268.14 tok/s | 5 GB peak94 W52°C2.85 tok/WQ4_K_M | ✓ Measured |
| Qwen2.5-Coder 14B | 149.75 tok/s | 8.6 GB peak117 W54°C1.28 tok/WQ4_K_M | ✓ Measured |
| Qwen3 32B | 71.15 tok/s | 19 GB peak83 W54°C0.86 tok/WQ4_K_M | ✓ Measured |
| Llama 3.3 70B | ✕ Won't fit needs ~46 GB | VRAM-gated at this precision | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Stable Diffusion XL | 21.08 images/min | 14.7 GB peak525 W57°C2.9 s/img | ✓ Measured |
| Z-Image Turbo | 10.73 images/min | 25.9 GB peak563 W64°C5.6 s/img | ✓ Measured |
| FLUX.1 dev | 2.04 images/min | 30.7 GB peak288 W68°C29.4 s/img | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Qwen-Image-Edit | ✕ Won't fit needs ~42 GB | VRAM-gated at this precision | ✓ Measured |
| Architecture | Blackwell (GB202) |
| CUDA cores | 21,760 |
| VRAM | 32GB GDDR7 |
| Memory bus | 512-bit |
| Memory bandwidth | 1792 GB/s |
| Boost clock | 2,407 MHz |
| TDP | 575 W |
| Process | 5nm |
| Interface | PCIe 5.0 x16 |
| Release date | 2025-01-30 |
| Launch MSRP | $1,999 |
NVIDIA GeForce RTX 5090 scores 22.3/100, #19 of 102. It ran 7 of 12; 2 exceeded its 32GB. Every figure here is our own measurement.
100% = this card, AI & Machine Learning headline metric (AI Score). #2 of 61 desktop cards in this vertical.
| GPU | Relative | % | AI Score |
|---|---|---|---|
| NVIDIA RTX 6000 Ada Generation | 107% | 23.8 | |
| NVIDIA GeForce RTX 5090 | 100% | 22.3 | |
| NVIDIA RTX 5880 Ada Generation | 73% | 16.3 | |
| NVIDIA RTX 5000 Ada Generation | 49% | 11 | |
| NVIDIA GeForce RTX 4090 | 47% | 10.4 | |
| NVIDIA GeForce RTX 3090 Ti | 38% | 8.5 |
Same card, other workloads: NVIDIA GeForce RTX 5090 Gaming benchmarks
← All AI & Machine Learning GPU rankings
| Transistors | 92,200 million |
| Die size | 750 mm² |
| Process node | 4 nm |
| Fabricated by | TSMC |
| Transistor density | 122.9 million per mm² |
Denser than 96% of the 76 cards we have silicon data for. Density is the clearest measure of what a process node bought: a card that gained it without growing the die got its speed from the fab rather than the architecture.
Silicon figures from Wikipedia (CC BY-SA 4.0). Benchmarks on this page are our own. Compare every chip.
Whole-job timings, composed from our measured per-model results on this card.
| Workflow | Time | Energy | Basis |
|---|---|---|---|
| Full codebase review | 6.7 min | 13.03 Wh | measured |
| 24-frame storyboard | 12.5 min | 21.46 Wh | all 2 stages measured |
Can't run: 20 long-form articles (needs Llama 3.3 70B).
This card is $1,999 to buy. The cheapest listed rate on Vast.ai is $0.321/hour, but that is the floor: we budget $0.385/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 5,190 GPU-hours. Below it you are paying for idle silicon.
| How you would use it | GPU-hours a year | Rental cost a year | Time to break even |
|---|---|---|---|
| 2 hours a day, hobby | 730 | $281 | 7.1 years |
| 8 hours a day, working on it | 2,920 | $1,125 | 1.8 years |
| 24/7, always-on agent | 8,760 | $3,374 | 7.1 months |
At hobby usage this card is very unlikely to pay for itself before it is superseded. Rent it. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.
Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.