16GB · AI Score 4.9/100 · first-party measured on 12 AI workloads
4.9 AI Score ✓ Measured
Every number on this page is first-party: GeForce RTX 5080 was run on our pinned 12-workload AI suite on 2026-07-11, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) GeForce RTX 5080 delivers about 150.71 tokens/sec. Llama 3.3 70B does not fit. It needs roughly 42GB and this card has 16GB. For image generation, SDXL runs at 4.43 it/s, while FLUX.1-dev won't fit at BF16 (needs ~26GB). 6 of the 12 workloads won't fit on 16GB at the tested precision, Qwen3 32B, Llama 3.3 70B, FLUX.1-dev, FLUX.1 Kontext and others. We publish those as hard gates rather than quietly dropping to a smaller quant.
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Qwen3 4B | 215.89 tok/s | 2.8 GB peak86 W53°C2.51 tok/WQ4_K_M | ✓ Measured |
| Llama 3.1 8B | 150.71 tok/s | 4.7 GB peak144 W56°C1.05 tok/WQ4_K_M | ✓ Measured |
| Qwen2.5-Coder 14B | 81.97 tok/s | 8.7 GB peak146 W59°C0.56 tok/WQ4_K_M | ✓ Measured |
| Qwen3 32B | ✕ Won't fit needs ~23 GB | VRAM-gated at this precision | ✓ Measured |
| Llama 3.3 70B | ✕ Won't fit needs ~46 GB | VRAM-gated at this precision | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Stable Diffusion XL | 8.86 images/min | 14.4 GB peak157 W62°C6.8 s/img | ✓ Measured |
| Z-Image Turbo | 2.7 images/min | 14.2 GB peak108 W60°C22.3 s/img | ✓ Measured |
| FLUX.1 dev | ✕ Won't fit needs ~26 GB | VRAM-gated at this precision | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| FLUX.1 Kontext dev | ✕ Won't fit needs ~26 GB | VRAM-gated at this precision | ✓ Measured |
| Qwen-Image-Edit | ✕ Won't fit needs ~42 GB | VRAM-gated at this precision | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| LTX-Video (distilled) | 3.34 frames/s | 9.3 GB peak120 W60°C29 s/clip | ✓ Measured |
| Wan 2.2 5B (720p) | ✕ Won't fit needs ~18 GB | VRAM-gated at this precision | ✓ Measured |
| Architecture | Blackwell (GB203) |
| CUDA cores | 10,752 |
| VRAM | 16GB GDDR7 |
| Memory bus | 256-bit |
| Memory bandwidth | 960 GB/s |
| Boost clock | 2,617 MHz |
| TDP | 360 W |
| Process | 5nm |
| Interface | PCIe 5.0 x16 |
| Release date | 2025-01-30 |
| Launch MSRP | $999 |
GeForce RTX 5080 scores 4.9/100, #47 of 102. It ran 6 of 12; 6 exceeded its 16GB. Every figure here is our own measurement.
100% = this card, AI & Machine Learning headline metric (AI Score). #12 of 61 desktop cards in this vertical.
| GPU | Relative | % | AI Score |
|---|---|---|---|
| NVIDIA GeForce RTX 3090 | 157% | 7.7 | |
| AMD Radeon RX 7900 XTX | 137% | 6.7 | |
| AMD Radeon RX 7900 XT | 127% | 6.2 | |
| GeForce RTX 4080 Super | 110% | 5.4 | |
| GeForce RTX 5080 | 100% | 4.9 | |
| NVIDIA GeForce RTX 4080 | 96% | 4.7 | |
| GeForce RTX 5070 Ti | 96% | 4.7 | |
| NVIDIA RTX 4000 (Ada Generation) | 90% | 4.4 | |
| NVIDIA GeForce RTX 4070 Ti Super | 88% | 4.3 |
Same card, other workloads: GeForce RTX 5080 Gaming benchmarks
← All AI & Machine Learning GPU rankings
| Transistors | 45,600 million |
| Die size | 378 mm² |
| Process node | 4 nm |
| Fabricated by | TSMC |
| Transistor density | 120.6 million per mm² |
Denser than 78% of the 76 cards we have silicon data for. Density is the clearest measure of what a process node bought: a card that gained it without growing the die got its speed from the fab rather than the architecture.
Silicon figures from Wikipedia (CC BY-SA 4.0). Benchmarks on this page are our own. Compare every chip.
Whole-job timings, composed from our measured per-model results on this card.
| Workflow | Time | Energy | Basis |
|---|---|---|---|
| Full codebase review | 12.2 min | 29.75 Wh | measured |
Can't run: 60-second AI short film (needs Qwen3 32B), 60-second AI short film, narrated (needs Qwen3 32B), 10 short social clips (needs Qwen3 32B), 40-product photo shoot (needs FLUX.1 Kontext dev), 6-panel comic page (needs Qwen3 32B), 20 long-form articles (needs Llama 3.3 70B), Character sheet, 12 poses (needs FLUX.1 dev), 100-photo restoration batch (needs FLUX.1 Kontext dev), 24-frame storyboard (needs Qwen3 32B), 100-photo restore and enlarge (needs FLUX.1 Kontext dev).
This card is $999 to buy. The cheapest listed rate on Vast.ai is $0.135/hour, but that is the floor: we budget $0.162/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 6,167 GPU-hours. Below it you are paying for idle silicon.
| How you would use it | GPU-hours a year | Rental cost a year | Time to break even |
|---|---|---|---|
| 2 hours a day, hobby | 730 | $118 | 8.4 years |
| 8 hours a day, working on it | 2,920 | $473 | 2.1 years |
| 24/7, always-on agent | 8,760 | $1,419 | 8.4 months |
At hobby usage this card is very unlikely to pay for itself before it is superseded. Rent it. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.
Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.