48GB · AI Score 8.2/100 · first-party measured on 12 AI workloads
8.2 AI Score Includes estimates
Every number on this page is first-party: NVIDIA Quadro RTX 8000 was run on our pinned 12-workload AI suite on 2026-07-11, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) NVIDIA Quadro RTX 8000 delivers about 84.68 tokens/sec. Stepping up to Qwen3 32B it holds roughly 21.98 tok/s. Llama 3.3 70B does not fit. It needs roughly 42GB and this card has 48GB. For image generation, SDXL runs at 3.09 it/s, and FLUX.1-dev at 0.08 it/s. 1 of the 12 workloads won't fit on 48GB at the tested precision, Llama 3.3 70B. We publish those as hard gates rather than quietly dropping to a smaller quant.
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| MiniCPM5 2B | 181.19 tok/s | 131 W36°CQ4_K_M | ✓ Measured |
| LFM2.5 2.6B | 192.62 tok/s | 121 W36°CQ4_K_M | ✓ Measured |
| Granite 4.1 3B | 138.18 tok/s | 148 W38°CQ4_K_M | ✓ Measured |
| Nemotron 3 Nano 4B | 134.78 tok/s | 153 W39°CQ4_K_M | ✓ Measured |
| Qwen3 4B | 127.01 tok/s | 2.9 GB peak113 W38°C1.13 tok/WQ4_K_M | ✓ Measured |
| DeepSeek Coder 7B Instruct v1.5 | 95.87 tok/s | 161 W40°CQ4_K_M | ✓ Measured |
| Llama 3 8B | 84.49 tok/s | 160 W40°CQ4_K_M | ✓ Measured |
| Llama 3.1 8B | 84.68 tok/s | 4.8 GB peak136 W44°C0.63 tok/WQ4_K_M | ✓ Measured |
| Qwen3 8B | 81.7 tok/s | 172 W40°CQ4_K_M | ✓ Measured |
| Nemotron Nano 9B v2 | 64.28 tok/s | 169 W41°CQ4_K_M | ✓ Measured |
| Ornith 1.5 9B | 72.54 tok/s | 158 W41°CQ4_K_M | ✓ Measured |
| Gemma 4 12B | 51.52 tok/s | 166 W42°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder 14B | 46.74 tok/s | 8.5 GB peak138 W49°C0.34 tok/WQ4_K_M | ✓ Measured |
| Qwen3 14B | 47.81 tok/s | 178 W42°CQ4_K_M | ✓ Measured |
| Gemma 4 26B A4B | 96.49 tok/s | 115 W38°CQ4_K_M | ✓ Measured |
| Qwen3 30B A3B | 127.94 tok/s | 112 W38°CQ4_K_M | ✓ Measured |
| Qwen3 32B | 21.98 tok/s | 18.6 GB peak125 W52°C0.18 tok/WQ4_K_M | ✓ Measured |
| Llama 3.3 70B | n/a tok/s | estimated | Est. |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Stable Diffusion XL | 6.18 images/min | 12.7 GB peak248 W53°C9.7 s/img | ✓ Measured |
| Z-Image Turbo | 0.45 images/min | 41.8 GB peak248 W60°C137.1 s/img | ✓ Measured |
| FLUX.1 dev | 0.17 images/min | 43.7 GB peak242 W60°C339.2 s/img | ✓ Measured |
| Architecture | Turing (TU102) |
| CUDA cores | 4,608 |
| VRAM | 48GB GDDR6 |
| Memory bus | 384-bit |
| Memory bandwidth | 672 GB/s |
| Boost clock | 1,770 MHz |
| TDP | 260 W |
| Process | 12nm |
| Interface | PCIe 3.0 x16 |
| Release date | 2018-08-27 |
| Launch MSRP | $9,999 |
NVIDIA Quadro RTX 8000 scores 8.2/100, #31 of 102. It ran 7 of 12; 1 exceeded its 48GB. Every figure here is our own measurement.
100% = this card, AI & Machine Learning headline metric (AI Score). #7 of 20 workstation cards in this vertical.
| GPU | Relative | % | AI Score |
|---|---|---|---|
| NVIDIA RTX PRO 5000 Blackwell | 385% | 31.6 | |
| NVIDIA RTX A6000 | 256% | 21 | |
| NVIDIA RTX PRO 4500 Blackwell | 157% | 12.9 | |
| NVIDIA RTX A5500 | 102% | 8.4 | |
| NVIDIA Quadro RTX 8000 | 100% | 8.2 | |
| NVIDIA RTX PRO 4000 Blackwell | 93% | 7.6 | |
| NVIDIA Quadro RTX 6000 (Turing) | 90% | 7.4 | |
| NVIDIA RTX A5000 | 90% | 7.4 | |
| NVIDIA RTX 4500 Ada Generation | 72% | 5.9 |
← All AI & Machine Learning GPU rankings
| Transistors | 18,600 million |
| Die size | 754 mm² |
| Process node | 12 nm |
| Fabricated by | TSMC |
| Transistor density | 24.7 million per mm² |
Denser than 69% of the 746 cards we have silicon data for. Density is the clearest measure of what a process node bought: a card that gained it without growing the die got its speed from the fab rather than the architecture.
Silicon figures from Wikipedia (CC BY-SA 4.0). Benchmarks on this page are our own. Compare every chip.
Whole-job timings, composed from our measured per-model results on this card.
| Workflow | Time | Energy | Basis |
|---|---|---|---|
| Full codebase review | 21.4 min | 49.24 Wh | measured |
| 24-frame storyboard | 54.9 min | 223 Wh | all 2 stages measured |
This card is $9,999 to buy. The cheapest listed rate on Vast.ai is $0.255/hour, but that is the floor: we budget $0.306/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 32,676 GPU-hours. Below it you are paying for idle silicon.
| How you would use it | GPU-hours a year | Rental cost a year | Time to break even |
|---|---|---|---|
| 2 hours a day, hobby | 730 | $223 | 44.8 years |
| 8 hours a day, working on it | 2,920 | $894 | 11.2 years |
| 24/7, always-on agent | 8,760 | $2,681 | 3.7 years |
At hobby usage this card is very unlikely to pay for itself before it is superseded. Rent it. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.
Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.