NVIDIA GeForce RTX 5090, AI & Machine Learning Benchmarks & Specs

32GB · AI Score 22.3/100 · first-party measured on 12 AI workloads

22.3 AI Score ✓ Measured

Every number on this page is first-party: NVIDIA GeForce RTX 5090 was run on our pinned 12-workload AI suite on 2026-07-11, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) NVIDIA GeForce RTX 5090 delivers about 268.14 tokens/sec. Stepping up to Qwen3 32B it holds roughly 71.15 tok/s. Llama 3.3 70B does not fit. It needs roughly 42GB and this card has 32GB. For image generation, SDXL runs at 10.54 it/s, and FLUX.1-dev at 0.95 it/s. 2 of the 12 workloads won't fit on 32GB at the tested precision, Llama 3.3 70B, Qwen-Image-Edit. We publish those as hard gates rather than quietly dropping to a smaller quant. NVIDIA GeForce RTX 5090 isn't a retail purchase for most people. It's rented by the hour. You can run this exact card on RunPod.

AI & Machine Learning benchmark results

Text Generation tok/s 37

Llama 3.2 1B1060.14
Qwen3 0.6B868.61
LFM2.5 2.6B566.57
MiniCPM5 2B507.31
gpt-oss-20b415.12
Nemotron 3 Nano 4B410.94
Nemotron 3.5 Lightning 30B A3B376.06
Qwen3 4B375.6
Granite 4.1 3B375.37
Qwen3-Coder 30B A3B365.84
Qwen3 30B A3B Instruct 2507352.33
Spark-X2.5-4B352.29
WorkloadResultTelemetryData
Qwen3 0.6B868.61 tok/s
118 W51°CQ4_K_M
✓ Measured
Llama 3.2 1B1060.14 tok/s
119 W51°CQ4_K_M
✓ Measured
MiniCPM5 2B507.31 tok/s
175 W51°CQ4_K_M
✓ Measured
LFM2.5 2.6B566.57 tok/s
142 W53°CQ4_K_M
✓ Measured
Granite 4.1 3B375.37 tok/s
190 W58°CQ4_K_M
✓ Measured
Agents-A1-4B328.96 tok/s
245 W54°CQ4_K_M
✓ Measured
Nemotron 3 Nano 4B410.94 tok/s
219 W54°CQ4_K_M
✓ Measured
Qwen3 4B375.6 tok/s
3 GB peak87 W48°C4.32 tok/WQ4_K_M
✓ Measured
Spark-X2.5-4B352.29 tok/s
264 W58°CQ4_K_M
✓ Measured
DeepSeek Coder 7B Instruct v1.5291.38 tok/s
266 W62°CQ4_K_M
✓ Measured
OLMo 3 7B Instruct266.68 tok/s
314 W56°CQ4_K_M
✓ Measured
OLMo 3 7B Think266.72 tok/s
288 W55°CQ4_K_M
✓ Measured
Qwen2-7B-Instruct277.64 tok/s
290 W56°CQ4_K_M
✓ Measured
Qwen2.5-7B284.6 tok/s
301 W59°CQ4_K_M
✓ Measured
Qwen2.5-Coder 7B284.59 tok/s
265 W60°CQ4_K_M
✓ Measured
Apertus-8B-Instruct257.69 tok/s
319 W55°CQ4_K_M
✓ Measured
Llama 3 8B259.18 tok/s
297 W56°CQ4_K_M
✓ Measured
Llama 3.1 8B268.14 tok/s
5 GB peak94 W52°C2.85 tok/WQ4_K_M
✓ Measured
Qwen3 8B243.91 tok/s
265 W45°CQ4_K_M
✓ Measured
Nemotron Nano 9B v2202.68 tok/s
291 W58°CQ4_K_M
✓ Measured
Ornith 1.5 9B225.79 tok/s
305 W56°CQ4_K_M
✓ Measured
Gemma 4 12B151.54 tok/s
278 W45°CQ4_K_M
✓ Measured
Qwen2.5-Coder 14B149.75 tok/s
8.6 GB peak117 W54°C1.28 tok/WQ4_K_M
✓ Measured
Qwen3 14B143.11 tok/s
314 W46°CQ4_K_M
✓ Measured
gpt-oss-20b415.12 tok/s
203 W53°CQ4_K_M
✓ Measured
Gemma 4 26B A4B264.69 tok/s
215 W52°CQ4_K_M
✓ Measured
Qwen3.6 27B76.6 tok/s
426 W57°CQ4_K_M
✓ Measured
Qwen3.8 27B75.74 tok/s
413 W57°CQ4_K_M
✓ Measured
Nemotron 3.5 Lightning 30B A3B376.06 tok/s
178 W55°CQ4_K_M
✓ Measured
Qwen3 30B A3B344.67 tok/s
156 W39°CQ4_K_M
✓ Measured
Qwen3 30B A3B Instruct 2507352.33 tok/s
189 W56°CQ4_K_M
✓ Measured
Qwen3-Coder 30B A3B365.84 tok/s
197 W54°CQ4_K_M
✓ Measured
Gemma 4 31B71.42 tok/s
439 W59°CQ4_K_M
✓ Measured
Qwen3 32B71.15 tok/s
19 GB peak83 W54°C0.86 tok/WQ4_K_M
✓ Measured
Ornith 1.5 35B A3B296.44 tok/s
188 W52°CQ4_K_M
✓ Measured
Qwen3.6 35B A3B300.13 tok/s
210 W54°CQ4_K_M
✓ Measured
Llama 3.3 70B✕ Won't fit needs ~46 GBVRAM-gated at this precision✓ Measured

Image Generation images/min 17

SDXL Turbo902.04
DreamShaper XL Lightning122.85
Stable Diffusion 1.5102.36
DreamShaper XL Turbo70.72
Stable Diffusion 2.165.53
FLUX.2 klein 4B59.33
Sana 1.6B51.52
PixArt-Sigma XL30.01
Stable Diffusion XL21.08
SSD-1B20.49
Kolors13.53
Playground v2.513.28
WorkloadResultTelemetryData
Stable Diffusion 1.5102.36 images/min
451 W57°C
✓ Measured
Stable Diffusion 2.165.53 images/min
309 W47°C
✓ Measured
SDXL Turbo902.04 images/min
33 W51°C
✓ Measured
SSD-1B20.49 images/min
589 W66°C
✓ Measured
Sana 1.6B51.52 images/min
549 W64°C
✓ Measured
Stable Diffusion XL21.08 images/min
14.7 GB peak525 W57°C2.9 s/img
✓ Measured
DreamShaper XL Lightning122.85 images/min
415 W62°C
✓ Measured
DreamShaper XL Turbo70.72 images/min
535 W50°C
✓ Measured
Playground v2.513.28 images/min
523 W63°C
✓ Measured
PixArt-Sigma XL30.01 images/min
547 W64°C
✓ Measured
FLUX.2 klein 4B59.33 images/min
518 W60°C
✓ Measured
Kolors13.53 images/min
593 W65°C
✓ Measured
Z-Image1.72 images/min
525 W75°C
✓ Measured
AuraFlow v0.34.35 images/min
572 W74°C
✓ Measured
Z-Image Turbo10.73 images/min
25.9 GB peak563 W64°C5.6 s/img
✓ Measured
Chroma1-HD1.29 images/min
521 W68°C
✓ Measured
FLUX.1 dev2.04 images/min
30.7 GB peak288 W68°C29.4 s/img
✓ Measured
CPU offload

Image Editing images/min 1

WorkloadResultTelemetryData
Qwen-Image-Edit✕ Won't fit needs ~42 GBVRAM-gated at this precision✓ Measured

Image to Video clips/min 5

LTX-Video (image to video)6.152
Stable Video Diffusion2.836
Wan 2.2 TI2V-5B (image to video)2.284
Stable Video Diffusion XT1.344
CogVideoX-5B I2V0.447
WorkloadResultTelemetryData
Stable Video Diffusion2.84 clips/min
562 W79°C
✓ Measured
LTX-Video (image to video)6.15 clips/min
530 W75°C
✓ Measured
Wan 2.2 TI2V-5B (image to video)2.28 clips/min
564 W70°C
✓ Measured
2 hosts ±0%
Stable Video Diffusion XT1.34 clips/min✓ Measured
2 hosts ±5% · CPU offload
CogVideoX-5B I2V0.45 clips/min✓ Measured
2 hosts ±1% · CPU offload

Video Generation frames/s 5

LTX-Video (distilled)10.449
Wan 2.1 1.3B0.879
CogVideoX-2B0.7
Wan 2.2 5B (720p)0.588
CogVideoX-5B0.259
WorkloadResultTelemetryData
Wan 2.1 1.3B0.88 frames/s
566 W75°C52.5 s/clip
✓ Measured
2 hosts ±6%
CogVideoX-2B0.7 frames/s
70 s/clip
✓ Measured
2 hosts ±3%
CogVideoX-5B0.26 frames/s
572 W78°C189 s/clip
✓ Measured
LTX-Video (distilled)10.45 frames/s
448 W50°C9.3 s/clip
✓ Measured
Wan 2.2 5B (720p)0.59 frames/s
456 W69°C83.4 s/clip
✓ Measured
CPU offload
How we measured this. Every result comes from our own pinned, reproducible AI suite, 12 workloads: the Qwen3-4B to Llama-70B LLM ladder (llama.cpp, Q4_K_M), SDXL / Z-Image / FLUX-dev generation, FLUX-Kontext / Qwen-Edit editing, and LTX / Wan video, run first-party on rented hardware with under 0.5% run-to-run variance. Peak VRAM, power draw, temperature and tokens-per-watt are captured per workload. “Won’t fit” rows are real data: where a model exceeds the card’s VRAM at the tested precision we record a hard gate rather than silently dropping to a smaller quant. Measured 2026-07-11 · harness 2.0.0-standalone.

NVIDIA GeForce RTX 5090 specifications

ArchitectureBlackwell (GB202)
CUDA cores21,760
VRAM32GB GDDR7
Memory bus512-bit
Memory bandwidth1792 GB/s
Boost clock2,407 MHz
TDP575 W
Process5nm
InterfacePCIe 5.0 x16
Release date2025-01-30
Launch MSRP$1,999

Verdict, capable, but 32GB sets the ceiling

NVIDIA GeForce RTX 5090 scores 22.3/100, #19 of 102. It ran 7 of 12; 2 exceeded its 32GB. Every figure here is our own measurement.

Relative performance: where the NVIDIA GeForce RTX 5090 lands

100% = this card, AI & Machine Learning headline metric (AI Score). #2 of 61 desktop cards in this vertical.

GPURelative%AI Score
NVIDIA RTX 6000 Ada Generation
107%23.8
NVIDIA GeForce RTX 5090
100%22.3
NVIDIA RTX 5880 Ada Generation
66%14.7
NVIDIA RTX 5000 Ada Generation
49%11
NVIDIA GeForce RTX 4090
48%10.6
NVIDIA GeForce RTX 3090 Ti
38%8.5

Same card, other workloads: NVIDIA GeForce RTX 5090 Gaming benchmarks

← All AI & Machine Learning GPU rankings

The silicon

Transistors92,200 million
Die size750 mm²
Process node4 nm
Fabricated byTSMC
Transistor density122.9 million per mm²

Denser than 96% of the 746 cards we have silicon data for. Density is the clearest measure of what a process node bought: a card that gained it without growing the die got its speed from the fab rather than the architecture.

Silicon figures from Wikipedia (CC BY-SA 4.0). Benchmarks on this page are our own. Compare every chip.

What this card can build

Whole-job timings, composed from our measured per-model results on this card.

WorkflowTimeEnergyBasis
Full codebase review6.7 min13.03 Whmeasured
Animate a batch of images8.8 min82.24 Whmeasured
60-second AI short film9.3 min23.85 Whall 3 stages measured
24-frame storyboard12.5 min21.46 Whall 2 stages measured
Short social clips24.8 min114.49 Whall 3 stages measured

Can't run: Long-form article batch (needs Llama 3.3 70B).

Rent or buy?

This card is $1,999 to buy. The cheapest listed rate on Vast.ai is $0.463/hour, but that is the floor: we budget $0.556/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 3,598 GPU-hours. Below it you are paying for idle silicon.

How you would use itGPU-hours a yearRental cost a yearTime to break even
2 hours a day, hobby730$4064.9 years
8 hours a day, working on it2,920$1,6221.2 years
24/7, always-on agent8,760$4,8674.9 months

At steady usage this card pays for itself inside a normal ownership window. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.

Rental price

$0.463/hr+72.1% since 2026-08-14low $0.213 · high $0.482

Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.