GeForce RTX 5080, AI & Machine Learning Benchmarks & Specs

16GB · AI Score 4.9/100 · first-party measured on 12 AI workloads

4.9 AI Score ✓ Measured

Every number on this page is first-party: GeForce RTX 5080 was run on our pinned 12-workload AI suite on 2026-07-11, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) GeForce RTX 5080 delivers about 150.71 tokens/sec. Llama 3.3 70B does not fit. It needs roughly 42GB and this card has 16GB. For image generation, SDXL runs at 4.43 it/s, while FLUX.1-dev won't fit at BF16 (needs ~26GB). 6 of the 12 workloads won't fit on 16GB at the tested precision, Qwen3 32B, Llama 3.3 70B, FLUX.1-dev, FLUX.1 Kontext and others. We publish those as hard gates rather than quietly dropping to a smaller quant.

AI & Machine Learning benchmark results

Text Generation tok/s 5

Qwen3 4B215.89
Llama 3.1 8B150.71
Qwen2.5-Coder 14B81.97
WorkloadResultTelemetryData
Qwen3 4B215.89 tok/s
2.8 GB peak86 W53°C2.51 tok/WQ4_K_M
✓ Measured
Llama 3.1 8B150.71 tok/s
4.7 GB peak144 W56°C1.05 tok/WQ4_K_M
✓ Measured
Qwen2.5-Coder 14B81.97 tok/s
8.7 GB peak146 W59°C0.56 tok/WQ4_K_M
✓ Measured
Qwen3 32B✕ Won't fit needs ~23 GBVRAM-gated at this precision✓ Measured
Llama 3.3 70B✕ Won't fit needs ~46 GBVRAM-gated at this precision✓ Measured

Image Generation images/min 3

WorkloadResultTelemetryData
Stable Diffusion XL8.86 images/min
14.4 GB peak157 W62°C6.8 s/img
✓ Measured
Z-Image Turbo2.7 images/min
14.2 GB peak108 W60°C22.3 s/img
✓ Measured
FLUX.1 dev✕ Won't fit needs ~26 GBVRAM-gated at this precision✓ Measured

Image Editing images/min 2

WorkloadResultTelemetryData
FLUX.1 Kontext dev✕ Won't fit needs ~26 GBVRAM-gated at this precision✓ Measured
Qwen-Image-Edit✕ Won't fit needs ~42 GBVRAM-gated at this precision✓ Measured

Video Generation frames/s 2

WorkloadResultTelemetryData
LTX-Video (distilled)3.34 frames/s
9.3 GB peak120 W60°C29 s/clip
✓ Measured
Wan 2.2 5B (720p)✕ Won't fit needs ~18 GBVRAM-gated at this precision✓ Measured
How we measured this. Every result comes from our own pinned, reproducible AI suite, 12 workloads: the Qwen3-4B to Llama-70B LLM ladder (llama.cpp, Q4_K_M), SDXL / Z-Image / FLUX-dev generation, FLUX-Kontext / Qwen-Edit editing, and LTX / Wan video, run first-party on rented hardware with under 0.5% run-to-run variance. Peak VRAM, power draw, temperature and tokens-per-watt are captured per workload. “Won’t fit” rows are real data: where a model exceeds the card’s VRAM at the tested precision we record a hard gate rather than silently dropping to a smaller quant. Measured 2026-07-11 · harness 2.0.0-standalone.

GeForce RTX 5080 specifications

ArchitectureBlackwell (GB203)
CUDA cores10,752
VRAM16GB GDDR7
Memory bus256-bit
Memory bandwidth960 GB/s
Boost clock2,617 MHz
TDP360 W
Process5nm
InterfacePCIe 5.0 x16
Release date2025-01-30
Launch MSRP$999

Verdict, capable, but 16GB sets the ceiling

GeForce RTX 5080 scores 4.9/100, #47 of 102. It ran 6 of 12; 6 exceeded its 16GB. Every figure here is our own measurement.

Relative performance: where the GeForce RTX 5080 lands

100% = this card, AI & Machine Learning headline metric (AI Score). #12 of 61 desktop cards in this vertical.

GPURelative%AI Score
NVIDIA GeForce RTX 3090
157%7.7
AMD Radeon RX 7900 XTX
137%6.7
AMD Radeon RX 7900 XT
127%6.2
GeForce RTX 4080 Super
110%5.4
GeForce RTX 5080
100%4.9
NVIDIA GeForce RTX 4080
96%4.7
GeForce RTX 5070 Ti
96%4.7
NVIDIA RTX 4000 (Ada Generation)
90%4.4
NVIDIA GeForce RTX 4070 Ti Super
88%4.3

Same card, other workloads: GeForce RTX 5080 Gaming benchmarks

← All AI & Machine Learning GPU rankings

The silicon

Transistors45,600 million
Die size378 mm²
Process node4 nm
Fabricated byTSMC
Transistor density120.6 million per mm²

Denser than 78% of the 76 cards we have silicon data for. Density is the clearest measure of what a process node bought: a card that gained it without growing the die got its speed from the fab rather than the architecture.

Silicon figures from Wikipedia (CC BY-SA 4.0). Benchmarks on this page are our own. Compare every chip.

What this card can build

Whole-job timings, composed from our measured per-model results on this card.

WorkflowTimeEnergyBasis
Full codebase review12.2 min29.75 Whmeasured

Can't run: 60-second AI short film (needs Qwen3 32B), 60-second AI short film, narrated (needs Qwen3 32B), 10 short social clips (needs Qwen3 32B), 40-product photo shoot (needs FLUX.1 Kontext dev), 6-panel comic page (needs Qwen3 32B), 20 long-form articles (needs Llama 3.3 70B), Character sheet, 12 poses (needs FLUX.1 dev), 100-photo restoration batch (needs FLUX.1 Kontext dev), 24-frame storyboard (needs Qwen3 32B), 100-photo restore and enlarge (needs FLUX.1 Kontext dev).

Rent or buy?

This card is $999 to buy. The cheapest listed rate on Vast.ai is $0.135/hour, but that is the floor: we budget $0.162/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 6,167 GPU-hours. Below it you are paying for idle silicon.

How you would use itGPU-hours a yearRental cost a yearTime to break even
2 hours a day, hobby730$1188.4 years
8 hours a day, working on it2,920$4732.1 years
24/7, always-on agent8,760$1,4198.4 months

At hobby usage this card is very unlikely to pay for itself before it is superseded. Rent it. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.

Rental price

$0.135/hr-5.6% since 2026-08-14low $0.084 · high $0.143

Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.