NVIDIA GeForce RTX 5090, AI & Machine Learning Benchmarks & Specs

32GB · AI Score 22.3/100 · first-party measured on 12 AI workloads

22.3 AI Score ✓ Measured

Every number on this page is first-party: NVIDIA GeForce RTX 5090 was run on our pinned 12-workload AI suite on 2026-07-11, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) NVIDIA GeForce RTX 5090 delivers about 268.14 tokens/sec. Stepping up to Qwen3 32B it holds roughly 71.15 tok/s. Llama 3.3 70B does not fit. It needs roughly 42GB and this card has 32GB. For image generation, SDXL runs at 10.54 it/s, and FLUX.1-dev at 0.95 it/s. 2 of the 12 workloads won't fit on 32GB at the tested precision, Llama 3.3 70B, Qwen-Image-Edit. We publish those as hard gates rather than quietly dropping to a smaller quant. NVIDIA GeForce RTX 5090 isn't a retail purchase for most people. It's rented by the hour. You can run this exact card on RunPod.

AI & Machine Learning benchmark results

Text Generation tok/s 5

Qwen3 4B375.6
Llama 3.1 8B268.14
Qwen2.5-Coder 14B149.75
Qwen3 32B71.15
WorkloadResultTelemetryData
Qwen3 4B375.6 tok/s
3 GB peak87 W48°C4.32 tok/WQ4_K_M
✓ Measured
Llama 3.1 8B268.14 tok/s
5 GB peak94 W52°C2.85 tok/WQ4_K_M
✓ Measured
Qwen2.5-Coder 14B149.75 tok/s
8.6 GB peak117 W54°C1.28 tok/WQ4_K_M
✓ Measured
Qwen3 32B71.15 tok/s
19 GB peak83 W54°C0.86 tok/WQ4_K_M
✓ Measured
Llama 3.3 70B✕ Won't fit needs ~46 GBVRAM-gated at this precision✓ Measured

Image Generation images/min 3

Stable Diffusion XL21.08
Z-Image Turbo10.725
FLUX.1 dev2.036
WorkloadResultTelemetryData
Stable Diffusion XL21.08 images/min
14.7 GB peak525 W57°C2.9 s/img
✓ Measured
Z-Image Turbo10.73 images/min
25.9 GB peak563 W64°C5.6 s/img
✓ Measured
FLUX.1 dev2.04 images/min
30.7 GB peak288 W68°C29.4 s/img
✓ Measured

Image Editing images/min 1

WorkloadResultTelemetryData
Qwen-Image-Edit✕ Won't fit needs ~42 GBVRAM-gated at this precision✓ Measured
How we measured this. Every result comes from our own pinned, reproducible AI suite, 12 workloads: the Qwen3-4B to Llama-70B LLM ladder (llama.cpp, Q4_K_M), SDXL / Z-Image / FLUX-dev generation, FLUX-Kontext / Qwen-Edit editing, and LTX / Wan video, run first-party on rented hardware with under 0.5% run-to-run variance. Peak VRAM, power draw, temperature and tokens-per-watt are captured per workload. “Won’t fit” rows are real data: where a model exceeds the card’s VRAM at the tested precision we record a hard gate rather than silently dropping to a smaller quant. Measured 2026-07-11 · harness 2.0.0-standalone.

NVIDIA GeForce RTX 5090 specifications

ArchitectureBlackwell (GB202)
CUDA cores21,760
VRAM32GB GDDR7
Memory bus512-bit
Memory bandwidth1792 GB/s
Boost clock2,407 MHz
TDP575 W
Process5nm
InterfacePCIe 5.0 x16
Release date2025-01-30
Launch MSRP$1,999

Verdict, capable, but 32GB sets the ceiling

NVIDIA GeForce RTX 5090 scores 22.3/100, #19 of 102. It ran 7 of 12; 2 exceeded its 32GB. Every figure here is our own measurement.

Relative performance: where the NVIDIA GeForce RTX 5090 lands

100% = this card, AI & Machine Learning headline metric (AI Score). #2 of 61 desktop cards in this vertical.

GPURelative%AI Score
NVIDIA RTX 6000 Ada Generation
107%23.8
NVIDIA GeForce RTX 5090
100%22.3
NVIDIA RTX 5880 Ada Generation
73%16.3
NVIDIA RTX 5000 Ada Generation
49%11
NVIDIA GeForce RTX 4090
47%10.4
NVIDIA GeForce RTX 3090 Ti
38%8.5

Same card, other workloads: NVIDIA GeForce RTX 5090 Gaming benchmarks

← All AI & Machine Learning GPU rankings

The silicon

Transistors92,200 million
Die size750 mm²
Process node4 nm
Fabricated byTSMC
Transistor density122.9 million per mm²

Denser than 96% of the 76 cards we have silicon data for. Density is the clearest measure of what a process node bought: a card that gained it without growing the die got its speed from the fab rather than the architecture.

Silicon figures from Wikipedia (CC BY-SA 4.0). Benchmarks on this page are our own. Compare every chip.

What this card can build

Whole-job timings, composed from our measured per-model results on this card.

WorkflowTimeEnergyBasis
Full codebase review6.7 min13.03 Whmeasured
24-frame storyboard12.5 min21.46 Whall 2 stages measured

Can't run: 20 long-form articles (needs Llama 3.3 70B).

Rent or buy?

This card is $1,999 to buy. The cheapest listed rate on Vast.ai is $0.321/hour, but that is the floor: we budget $0.385/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 5,190 GPU-hours. Below it you are paying for idle silicon.

How you would use itGPU-hours a yearRental cost a yearTime to break even
2 hours a day, hobby730$2817.1 years
8 hours a day, working on it2,920$1,1251.8 years
24/7, always-on agent8,760$3,3747.1 months

At hobby usage this card is very unlikely to pay for itself before it is superseded. Rent it. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.

Rental price

$0.321/hr+19.3% since 2026-08-14low $0.269 · high $0.336

Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.