32GB · AI Score 22.3/100 · first-party measured on 12 AI workloads
22.3 AI Score ✓ Measured
Every number on this page is first-party: NVIDIA GeForce RTX 5090 was run on our pinned 12-workload AI suite on 2026-07-11, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) NVIDIA GeForce RTX 5090 delivers about 268.14 tokens/sec. Stepping up to Qwen3 32B it holds roughly 71.15 tok/s. Llama 3.3 70B does not fit. It needs roughly 42GB and this card has 32GB. For image generation, SDXL runs at 10.54 it/s, and FLUX.1-dev at 0.95 it/s. 2 of the 12 workloads won't fit on 32GB at the tested precision, Llama 3.3 70B, Qwen-Image-Edit. We publish those as hard gates rather than quietly dropping to a smaller quant. NVIDIA GeForce RTX 5090 isn't a retail purchase for most people. It's rented by the hour. You can run this exact card on RunPod.
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Qwen3 0.6B | 868.61 tok/s | 118 W51°CQ4_K_M | ✓ Measured |
| Llama 3.2 1B | 1060.14 tok/s | 119 W51°CQ4_K_M | ✓ Measured |
| MiniCPM5 2B | 507.31 tok/s | 175 W51°CQ4_K_M | ✓ Measured |
| LFM2.5 2.6B | 566.57 tok/s | 142 W53°CQ4_K_M | ✓ Measured |
| Granite 4.1 3B | 375.37 tok/s | 190 W58°CQ4_K_M | ✓ Measured |
| Agents-A1-4B | 328.96 tok/s | 245 W54°CQ4_K_M | ✓ Measured |
| Nemotron 3 Nano 4B | 410.94 tok/s | 219 W54°CQ4_K_M | ✓ Measured |
| Qwen3 4B | 375.6 tok/s | 3 GB peak87 W48°C4.32 tok/WQ4_K_M | ✓ Measured |
| Spark-X2.5-4B | 352.29 tok/s | 264 W58°CQ4_K_M | ✓ Measured |
| DeepSeek Coder 7B Instruct v1.5 | 291.38 tok/s | 266 W62°CQ4_K_M | ✓ Measured |
| OLMo 3 7B Instruct | 266.68 tok/s | 314 W56°CQ4_K_M | ✓ Measured |
| OLMo 3 7B Think | 266.72 tok/s | 288 W55°CQ4_K_M | ✓ Measured |
| Qwen2-7B-Instruct | 277.64 tok/s | 290 W56°CQ4_K_M | ✓ Measured |
| Qwen2.5-7B | 284.6 tok/s | 301 W59°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder 7B | 284.59 tok/s | 265 W60°CQ4_K_M | ✓ Measured |
| Apertus-8B-Instruct | 257.69 tok/s | 319 W55°CQ4_K_M | ✓ Measured |
| Llama 3 8B | 259.18 tok/s | 297 W56°CQ4_K_M | ✓ Measured |
| Llama 3.1 8B | 268.14 tok/s | 5 GB peak94 W52°C2.85 tok/WQ4_K_M | ✓ Measured |
| Qwen3 8B | 243.91 tok/s | 265 W45°CQ4_K_M | ✓ Measured |
| Nemotron Nano 9B v2 | 202.68 tok/s | 291 W58°CQ4_K_M | ✓ Measured |
| Ornith 1.5 9B | 225.79 tok/s | 305 W56°CQ4_K_M | ✓ Measured |
| Gemma 4 12B | 151.54 tok/s | 278 W45°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder 14B | 149.75 tok/s | 8.6 GB peak117 W54°C1.28 tok/WQ4_K_M | ✓ Measured |
| Qwen3 14B | 143.11 tok/s | 314 W46°CQ4_K_M | ✓ Measured |
| gpt-oss-20b | 415.12 tok/s | 203 W53°CQ4_K_M | ✓ Measured |
| Gemma 4 26B A4B | 264.69 tok/s | 215 W52°CQ4_K_M | ✓ Measured |
| Qwen3.6 27B | 76.6 tok/s | 426 W57°CQ4_K_M | ✓ Measured |
| Qwen3.8 27B | 75.74 tok/s | 413 W57°CQ4_K_M | ✓ Measured |
| Nemotron 3.5 Lightning 30B A3B | 376.06 tok/s | 178 W55°CQ4_K_M | ✓ Measured |
| Qwen3 30B A3B | 344.67 tok/s | 156 W39°CQ4_K_M | ✓ Measured |
| Qwen3 30B A3B Instruct 2507 | 352.33 tok/s | 189 W56°CQ4_K_M | ✓ Measured |
| Qwen3-Coder 30B A3B | 365.84 tok/s | 197 W54°CQ4_K_M | ✓ Measured |
| Gemma 4 31B | 71.42 tok/s | 439 W59°CQ4_K_M | ✓ Measured |
| Qwen3 32B | 71.15 tok/s | 19 GB peak83 W54°C0.86 tok/WQ4_K_M | ✓ Measured |
| Ornith 1.5 35B A3B | 296.44 tok/s | 188 W52°CQ4_K_M | ✓ Measured |
| Qwen3.6 35B A3B | 300.13 tok/s | 210 W54°CQ4_K_M | ✓ Measured |
| Llama 3.3 70B | ✕ Won't fit needs ~46 GB | VRAM-gated at this precision | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Stable Diffusion 1.5 | 102.36 images/min | 451 W57°C | ✓ Measured |
| Stable Diffusion 2.1 | 65.53 images/min | 309 W47°C | ✓ Measured |
| SDXL Turbo | 902.04 images/min | 33 W51°C | ✓ Measured |
| SSD-1B | 20.49 images/min | 589 W66°C | ✓ Measured |
| Sana 1.6B | 51.52 images/min | 549 W64°C | ✓ Measured |
| Stable Diffusion XL | 21.08 images/min | 14.7 GB peak525 W57°C2.9 s/img | ✓ Measured |
| DreamShaper XL Lightning | 122.85 images/min | 415 W62°C | ✓ Measured |
| DreamShaper XL Turbo | 70.72 images/min | 535 W50°C | ✓ Measured |
| Playground v2.5 | 13.28 images/min | 523 W63°C | ✓ Measured |
| PixArt-Sigma XL | 30.01 images/min | 547 W64°C | ✓ Measured |
| FLUX.2 klein 4B | 59.33 images/min | 518 W60°C | ✓ Measured |
| Kolors | 13.53 images/min | 593 W65°C | ✓ Measured |
| Z-Image | 1.72 images/min | 525 W75°C | ✓ Measured |
| AuraFlow v0.3 | 4.35 images/min | 572 W74°C | ✓ Measured |
| Z-Image Turbo | 10.73 images/min | 25.9 GB peak563 W64°C5.6 s/img | ✓ Measured |
| Chroma1-HD | 1.29 images/min | 521 W68°C | ✓ Measured |
| FLUX.1 dev | 2.04 images/min | 30.7 GB peak288 W68°C29.4 s/img | ✓ Measured CPU offload |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Qwen-Image-Edit | ✕ Won't fit needs ~42 GB | VRAM-gated at this precision | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Stable Video Diffusion | 2.84 clips/min | 562 W79°C | ✓ Measured |
| LTX-Video (image to video) | 6.15 clips/min | 530 W75°C | ✓ Measured |
| Wan 2.2 TI2V-5B (image to video) | 2.28 clips/min | 564 W70°C | ✓ Measured 2 hosts ±0% |
| Stable Video Diffusion XT | 1.34 clips/min | ✓ Measured 2 hosts ±5% · CPU offload | |
| CogVideoX-5B I2V | 0.45 clips/min | ✓ Measured 2 hosts ±1% · CPU offload |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Wan 2.1 1.3B | 0.88 frames/s | 566 W75°C52.5 s/clip | ✓ Measured 2 hosts ±6% |
| CogVideoX-2B | 0.7 frames/s | 70 s/clip | ✓ Measured 2 hosts ±3% |
| CogVideoX-5B | 0.26 frames/s | 572 W78°C189 s/clip | ✓ Measured |
| LTX-Video (distilled) | 10.45 frames/s | 448 W50°C9.3 s/clip | ✓ Measured |
| Wan 2.2 5B (720p) | 0.59 frames/s | 456 W69°C83.4 s/clip | ✓ Measured CPU offload |
| Architecture | Blackwell (GB202) |
| CUDA cores | 21,760 |
| VRAM | 32GB GDDR7 |
| Memory bus | 512-bit |
| Memory bandwidth | 1792 GB/s |
| Boost clock | 2,407 MHz |
| TDP | 575 W |
| Process | 5nm |
| Interface | PCIe 5.0 x16 |
| Release date | 2025-01-30 |
| Launch MSRP | $1,999 |
NVIDIA GeForce RTX 5090 scores 22.3/100, #19 of 102. It ran 7 of 12; 2 exceeded its 32GB. Every figure here is our own measurement.
100% = this card, AI & Machine Learning headline metric (AI Score). #2 of 61 desktop cards in this vertical.
| GPU | Relative | % | AI Score |
|---|---|---|---|
| NVIDIA RTX 6000 Ada Generation | 107% | 23.8 | |
| NVIDIA GeForce RTX 5090 | 100% | 22.3 | |
| NVIDIA RTX 5880 Ada Generation | 66% | 14.7 | |
| NVIDIA RTX 5000 Ada Generation | 49% | 11 | |
| NVIDIA GeForce RTX 4090 | 48% | 10.6 | |
| NVIDIA GeForce RTX 3090 Ti | 38% | 8.5 |
Same card, other workloads: NVIDIA GeForce RTX 5090 Gaming benchmarks
← All AI & Machine Learning GPU rankings
| Transistors | 92,200 million |
| Die size | 750 mm² |
| Process node | 4 nm |
| Fabricated by | TSMC |
| Transistor density | 122.9 million per mm² |
Denser than 96% of the 746 cards we have silicon data for. Density is the clearest measure of what a process node bought: a card that gained it without growing the die got its speed from the fab rather than the architecture.
Silicon figures from Wikipedia (CC BY-SA 4.0). Benchmarks on this page are our own. Compare every chip.
Whole-job timings, composed from our measured per-model results on this card.
| Workflow | Time | Energy | Basis |
|---|---|---|---|
| Full codebase review | 6.7 min | 13.03 Wh | measured |
| Animate a batch of images | 8.8 min | 82.24 Wh | measured |
| 60-second AI short film | 9.3 min | 23.85 Wh | all 3 stages measured |
| 24-frame storyboard | 12.5 min | 21.46 Wh | all 2 stages measured |
| Short social clips | 24.8 min | 114.49 Wh | all 3 stages measured |
Can't run: Long-form article batch (needs Llama 3.3 70B).
This card is $1,999 to buy. The cheapest listed rate on Vast.ai is $0.463/hour, but that is the floor: we budget $0.556/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 3,598 GPU-hours. Below it you are paying for idle silicon.
| How you would use it | GPU-hours a year | Rental cost a year | Time to break even |
|---|---|---|---|
| 2 hours a day, hobby | 730 | $406 | 4.9 years |
| 8 hours a day, working on it | 2,920 | $1,622 | 1.2 years |
| 24/7, always-on agent | 8,760 | $4,867 | 4.9 months |
At steady usage this card pays for itself inside a normal ownership window. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.
Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.