NVIDIA GeForce RTX 4090, AI & Machine Learning Benchmarks & Specs

24GB · AI Score 10.4/100 · first-party measured on 12 AI workloads

10.6 AI Score ✓ Measured

Every number on this page is first-party: NVIDIA GeForce RTX 4090 was run on our pinned 12-workload AI suite on 2026-07-11, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) NVIDIA GeForce RTX 4090 delivers about 171.29 tokens/sec. Stepping up to Qwen3 32B it holds roughly 44.28 tok/s. Llama 3.3 70B does not fit. It needs roughly 42GB and this card has 24GB. For image generation, SDXL runs at 8.14 it/s, while FLUX.1-dev won't fit at BF16 (needs ~26GB). 4 of the 12 workloads won't fit on 24GB at the tested precision, Llama 3.3 70B, FLUX.1-dev, FLUX.1 Kontext, Qwen-Image-Edit. We publish those as hard gates rather than quietly dropping to a smaller quant. NVIDIA GeForce RTX 4090 isn't a retail purchase for most people. It's rented by the hour. You can run this exact card on RunPod.

Bench notes: from the person who ran it

Best consumer efficiency I measured, period: 2.24 tokens/watt on Qwen3 4B. It'll take its full 450W on SDXL and stay under 68°C. The one thing people get wrong: 24GB still can't hold Llama 3.3 70B: that model wants ~46GB, and no consumer card changes that. Quick note on the setup: all my AI benchmarking was done on rented cloud GPUs, I used all three of Vast.ai, RunPod and Modal depending on which had the card, and they all have their pros and cons. Same pinned harness on every run, and everything here got double-checked before it went up.

AI & Machine Learning benchmark results

Text Generation tok/s 34

Qwen3 0.6B770.13
Llama 3.2 1B752.52
LFM2.5 2.6B385.05
MiniCPM5 2B372.72
gpt-oss-20b286.23
Granite 4.1 3B276.62
Qwen3-Coder 30B A3B271.04
Qwen3 30B A3B Instruct 2507263.17
Qwen3 4B260.54
Qwen3 30B A3B259.61
Nemotron 3 Nano 4B258.12
Spark-X2.5-4B245.97
WorkloadResultTelemetryData
Qwen3 0.6B770.13 tok/s
86 W36°CQ4_K_M
✓ Measured
Llama 3.2 1B752.52 tok/s
95 W42°CQ4_K_M
✓ Measured
MiniCPM5 2B372.72 tok/s
105 W56°CQ4_K_M
✓ Measured
LFM2.5 2.6B385.05 tok/s
142 W57°CQ4_K_M
✓ Measured
Granite 4.1 3B276.62 tok/s
143 W60°CQ4_K_M
✓ Measured
Agents-A1-4B219.4 tok/s
160 W33°CQ4_K_M
✓ Measured
Nemotron 3 Nano 4B258.12 tok/s
166 W56°CQ4_K_M
✓ Measured
Qwen3 4B260.54 tok/s
3.1 GB peak116 W38°C2.24 tok/WQ4_K_M
✓ Measured
Spark-X2.5-4B245.97 tok/s
166 W33°CQ4_K_M
✓ Measured
DeepSeek Coder 7B Instruct v1.5188.59 tok/s
179 W63°CQ4_K_M
✓ Measured
OLMo 3 7B Instruct170.75 tok/s
186 W34°CQ4_K_M
✓ Measured
OLMo 3 7B Think170.69 tok/s
183 W34°CQ4_K_M
✓ Measured
Qwen2-7B-Instruct177.4 tok/s
189 W35°CQ4_K_M
✓ Measured
Qwen2.5-7B183.66 tok/s
221 W47°CQ4_K_M
✓ Measured
Qwen2.5-Coder 7B183.73 tok/s
221 W49°CQ4_K_M
✓ Measured
Apertus-8B-Instruct163.09 tok/s
191 W36°CQ4_K_M
✓ Measured
Llama 3 8B167.69 tok/s
199 W54°CQ4_K_M
✓ Measured
Llama 3.1 8B171.29 tok/s
5.1 GB peak191 W42°C0.9 tok/WQ4_K_M
✓ Measured
Qwen3 8B164.32 tok/s
214 W58°CQ4_K_M
✓ Measured
Nemotron Nano 9B v2121.89 tok/s
206 W55°CQ4_K_M
✓ Measured
Ornith 1.5 9B144.04 tok/s
205 W57°CQ4_K_M
✓ Measured
Gemma 4 12B103.57 tok/s
219 W59°CQ4_K_M
✓ Measured
Qwen2.5-Coder 14B95.12 tok/s
8.5 GB peak174 W46°C0.55 tok/WQ4_K_M
✓ Measured
Qwen3 14B96.37 tok/s
238 W60°CQ4_K_M
✓ Measured
gpt-oss-20b286.23 tok/s
154 W47°CQ4_K_M
✓ Measured
Gemma 4 26B A4B188.23 tok/s
139 W55°CQ4_K_M
✓ Measured
Qwen3.6 27B49.02 tok/s
261 W65°CQ4_K_M
✓ Measured
Qwen3.8 27B48.01 tok/s
259 W63°CQ4_K_M
✓ Measured
Qwen3 30B A3B259.61 tok/s
135 W56°CQ4_K_M
✓ Measured
Qwen3 30B A3B Instruct 2507263.17 tok/s
128 W59°CQ4_K_M
✓ Measured
Qwen3-Coder 30B A3B271.04 tok/s
146 W45°CQ4_K_M
✓ Measured
Gemma 4 31B45.08 tok/s
266 W65°CQ4_K_M
✓ Measured
Qwen3 32B44.28 tok/s
18.9 GB peak164 W47°C0.27 tok/WQ4_K_M
✓ Measured
Llama 3.3 70B✕ Won't fit needs ~46 GBVRAM-gated at this precision✓ Measured

Image Generation images/min 20

SD Turbo786.15
SDXL Turbo652.11
LCM DreamShaper v7260.94
DreamShaper XL Lightning91.52
SDXL-Lightning86.705
Stable Diffusion 1.580.14
Stable Diffusion 2.160.58
DreamShaper XL Turbo54.35
FLUX.2 klein 4B42.38
Sana 1.6B41.47
PixArt-Sigma XL24.27
Stable Diffusion XL16.28
WorkloadResultTelemetryData
Stable Diffusion 1.580.14 images/min
325 W53°C
✓ Measured
SD Turbo786.15 images/min
91 W50°C
✓ Measured
Stable Diffusion 2.160.58 images/min
275 W38°C
✓ Measured
LCM DreamShaper v7260.94 images/min
178 W60°C
✓ Measured
SDXL Turbo652.11 images/min
90 W34°C
✓ Measured
SSD-1B15.77 images/min
414 W68°C
✓ Measured
SDXL-Lightning86.71 images/min
271 W49°C
✓ Measured
Z-Image Turbo7.24 images/min
414 W78°C
✓ Measured
Sana 1.6B41.47 images/min
411 W59°C
✓ Measured
Stable Diffusion XL16.28 images/min
14.8 GB peak421 W53°C3.7 s/img
✓ Measured
DreamShaper XL Lightning91.52 images/min
306 W71°C
✓ Measured
DreamShaper XL Turbo54.35 images/min
412 W74°C
✓ Measured
Playground v2.59.96 images/min
396 W52°C
✓ Measured
PixArt-Sigma XL24.27 images/min
415 W60°C
✓ Measured
Stable Diffusion 3 Medium13.95 images/min
437 W73°C
✓ Measured
FLUX.2 klein 4B42.38 images/min
392 W50°C
✓ Measured
Kolors9.88 images/min
417 W70°C
✓ Measured
Z-Image1.17 images/min
398 W59°C
✓ Measured
AuraFlow v0.33.02 images/min
423 W61°C
✓ Measured
FLUX.1 dev✕ Won't fit needs ~26 GBVRAM-gated at this precision✓ Measured

Image Editing images/min 2

WorkloadResultTelemetryData
FLUX.1 Kontext dev✕ Won't fit needs ~26 GBVRAM-gated at this precision✓ Measured
Qwen-Image-Edit✕ Won't fit needs ~42 GBVRAM-gated at this precision✓ Measured

Image to Video clips/min 6

LTX-Video (image to video)4.188
Stable Video Diffusion2.235
Stable Video Diffusion XT1.212
Wan 2.2 TI2V-5B (image to video)1.122
CogVideoX-5B I2V0.328
Cosmos-Predict2 2B Video2World0.058
WorkloadResultTelemetryData
Stable Video Diffusion2.24 clips/min
403 W74°C
✓ Measured
LTX-Video (image to video)4.19 clips/min
384 W73°C
✓ Measured
Wan 2.2 TI2V-5B (image to video)1.12 clips/min✓ Measured
CPU offload
Cosmos-Predict2 2B Video2World0.06 clips/min
424 W66°C
✓ Measured
CPU offload
Stable Video Diffusion XT1.21 clips/min
347 W68°C
✓ Measured
2 hosts ±2% · CPU offload
CogVideoX-5B I2V0.33 clips/min
344 W70°C
✓ Measured
2 hosts ±1% · CPU offload

Video Generation frames/s 6

AnimateDiff-Lightning10.631
LTX-Video (distilled)6.696
Wan 2.1 1.3B0.673
CogVideoX-2B0.517
Wan 2.2 5B (720p)0.43
CogVideoX-5B0.17
WorkloadResultTelemetryData
Wan 2.1 1.3B0.67 frames/s
411 W73°C72.9 s/clip
✓ Measured
AnimateDiff-Lightning10.63 frames/s
389 W56°C
✓ Measured
CogVideoX-2B0.52 frames/s
410 W53°C96.5 s/clip
✓ Measured
2 hosts ±2%
CogVideoX-5B0.17 frames/s
423 W85°C287.9 s/clip
✓ Measured
CPU offload
LTX-Video (distilled)6.7 frames/s
227 W45°C14.5 s/clip
✓ Measured
3 hosts ±25%
Wan 2.2 5B (720p)0.43 frames/s
18.6 GB peak366 W68°C114.5 s/clip
✓ Measured
CPU offload

Image to 3D assets/hour 5

Stable Fast 3D9113.924
Shap-E (image to 3D)1309.091
TRELLIS.2 Image-to-3D57.2
Hunyuan3D 2mini19.813
Hunyuan3D 2.012.266
WorkloadResultTelemetryData
Shap-E (image to 3D)1309.09 assets/hour
362 W47°C
✓ Measured
Stable Fast 3D9113.92 assets/hour
157 W37°C
✓ Measured
Hunyuan3D 2mini19.81 assets/hour
71 W50°C
✓ Measured
TRELLIS.2 Image-to-3D57.2 assets/hour✓ Measured
Hunyuan3D 2.012.27 assets/hour
68 W54°C
✓ Measured
How we measured this. Every result comes from our own pinned, reproducible AI suite, 12 workloads: the Qwen3-4B to Llama-70B LLM ladder (llama.cpp, Q4_K_M), SDXL / Z-Image / FLUX-dev generation, FLUX-Kontext / Qwen-Edit editing, and LTX / Wan video, run first-party on rented hardware with under 0.5% run-to-run variance. Peak VRAM, power draw, temperature and tokens-per-watt are captured per workload. “Won’t fit” rows are real data: where a model exceeds the card’s VRAM at the tested precision we record a hard gate rather than silently dropping to a smaller quant. Measured 2026-07-11 · harness 2.0.0-standalone.

NVIDIA GeForce RTX 4090 specifications

ArchitectureAda Lovelace
CUDA cores16,384
VRAM24GB GDDR6X
Memory bus384-bit
Memory bandwidth1008 GB/s
Boost clock2,520 MHz
TDP450 W
Process4nm
InterfacePCIe 4.0 x16
Release date2022-10-12
Launch MSRP$1,599

Verdict, capable, but 24GB sets the ceiling

NVIDIA GeForce RTX 4090 scores 10.4/100, #28 of 102. It ran 8 of 12; 4 exceeded its 24GB. Every figure here is our own measurement.

Relative performance: where the NVIDIA GeForce RTX 4090 lands

100% = this card, AI & Machine Learning headline metric (AI Score). #5 of 61 desktop cards in this vertical.

GPURelative%AI Score
NVIDIA RTX 6000 Ada Generation
225%23.8
NVIDIA GeForce RTX 5090
210%22.3
NVIDIA RTX 5880 Ada Generation
139%14.7
NVIDIA RTX 5000 Ada Generation
104%11
NVIDIA GeForce RTX 4090
100%10.6
NVIDIA GeForce RTX 3090 Ti
80%8.5
NVIDIA Titan RTX
77%8.2
NVIDIA GeForce RTX 3090
73%7.7
GeForce RTX 5080
49%5.2

Same card, other workloads: NVIDIA GeForce RTX 4090 Gaming benchmarks

← All AI & Machine Learning GPU rankings

The silicon

Transistors76,300 million
Die size608.4 mm²
Process node4 nm
Fabricated byTSMC
Transistor density125.4 million per mm²

Denser than 97% of the 746 cards we have silicon data for. Density is the clearest measure of what a process node bought: a card that gained it without growing the die got its speed from the fab rather than the architecture.

Silicon figures from Wikipedia (CC BY-SA 4.0). Benchmarks on this page are our own. Compare every chip.

What this card can build

Whole-job timings, composed from our measured per-model results on this card.

WorkflowTimeEnergyBasis
24-frame storyboard3.8 min24.33 Whall 2 stages measured
60-second AI short film5.8 min21.08 Whall 3 stages measured
Full codebase review10.5 min30.47 Whmeasured
Animate a batch of images17.8 minn/ameasured
Short social clips21.5 min125.78 Whall 3 stages measured

Can't run: Product photo shoot (needs FLUX.1 Kontext dev), Photo restoration batch (needs FLUX.1 Kontext dev), Restore and enlarge photos (needs FLUX.1 Kontext dev), Character sheet, 12 poses (needs FLUX.1 dev), 6-panel comic page (needs FLUX.1 dev), Long-form article batch (needs Llama 3.3 70B).

Rent or buy?

This card is $1,599 to buy. The cheapest listed rate on RunPod is $0.340/hour, but that is the floor: we budget $0.408/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 3,919 GPU-hours. Below it you are paying for idle silicon.

How you would use itGPU-hours a yearRental cost a yearTime to break even
2 hours a day, hobby730$2985.4 years
8 hours a day, working on it2,920$1,1911.3 years
24/7, always-on agent8,760$3,5745.4 months

At hobby usage this card is very unlikely to pay for itself before it is superseded. Rent it. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.

Rental price

$0.340/hr+150.0% since 2026-08-14low $0.121 · high $0.340

Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.