NVIDIA Quadro RTX 8000, AI & Machine Learning Benchmarks & Specs

48GB · AI Score 8.2/100 · first-party measured on 12 AI workloads

8.2 AI Score Includes estimates

Every number on this page is first-party: NVIDIA Quadro RTX 8000 was run on our pinned 12-workload AI suite on 2026-07-11, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) NVIDIA Quadro RTX 8000 delivers about 84.68 tokens/sec. Stepping up to Qwen3 32B it holds roughly 21.98 tok/s. Llama 3.3 70B does not fit. It needs roughly 42GB and this card has 48GB. For image generation, SDXL runs at 3.09 it/s, and FLUX.1-dev at 0.08 it/s. 1 of the 12 workloads won't fit on 48GB at the tested precision, Llama 3.3 70B. We publish those as hard gates rather than quietly dropping to a smaller quant.

AI & Machine Learning benchmark results

Text Generation tok/s 18

LFM2.5 2.6B192.62
MiniCPM5 2B181.19
Granite 4.1 3B138.18
Nemotron 3 Nano 4B134.78
Qwen3 30B A3B127.94
Qwen3 4B127.01
Gemma 4 26B A4B96.49
DeepSeek Coder 7B Instruct v1.595.87
Llama 3.1 8B84.68
Llama 3 8B84.49
Qwen3 8B81.7
Ornith 1.5 9B72.54
WorkloadResultTelemetryData
MiniCPM5 2B181.19 tok/s
131 W36°CQ4_K_M
✓ Measured
LFM2.5 2.6B192.62 tok/s
121 W36°CQ4_K_M
✓ Measured
Granite 4.1 3B138.18 tok/s
148 W38°CQ4_K_M
✓ Measured
Nemotron 3 Nano 4B134.78 tok/s
153 W39°CQ4_K_M
✓ Measured
Qwen3 4B127.01 tok/s
2.9 GB peak113 W38°C1.13 tok/WQ4_K_M
✓ Measured
DeepSeek Coder 7B Instruct v1.595.87 tok/s
161 W40°CQ4_K_M
✓ Measured
Llama 3 8B84.49 tok/s
160 W40°CQ4_K_M
✓ Measured
Llama 3.1 8B84.68 tok/s
4.8 GB peak136 W44°C0.63 tok/WQ4_K_M
✓ Measured
Qwen3 8B81.7 tok/s
172 W40°CQ4_K_M
✓ Measured
Nemotron Nano 9B v264.28 tok/s
169 W41°CQ4_K_M
✓ Measured
Ornith 1.5 9B72.54 tok/s
158 W41°CQ4_K_M
✓ Measured
Gemma 4 12B51.52 tok/s
166 W42°CQ4_K_M
✓ Measured
Qwen2.5-Coder 14B46.74 tok/s
8.5 GB peak138 W49°C0.34 tok/WQ4_K_M
✓ Measured
Qwen3 14B47.81 tok/s
178 W42°CQ4_K_M
✓ Measured
Gemma 4 26B A4B96.49 tok/s
115 W38°CQ4_K_M
✓ Measured
Qwen3 30B A3B127.94 tok/s
112 W38°CQ4_K_M
✓ Measured
Qwen3 32B21.98 tok/s
18.6 GB peak125 W52°C0.18 tok/WQ4_K_M
✓ Measured
Llama 3.3 70Bn/a tok/sestimatedEst.

Image Generation images/min 3

Stable Diffusion XL6.18
Z-Image Turbo0.45
FLUX.1 dev0.171
WorkloadResultTelemetryData
Stable Diffusion XL6.18 images/min
12.7 GB peak248 W53°C9.7 s/img
✓ Measured
Z-Image Turbo0.45 images/min
41.8 GB peak248 W60°C137.1 s/img
✓ Measured
FLUX.1 dev0.17 images/min
43.7 GB peak242 W60°C339.2 s/img
✓ Measured
How we measured this. Every result comes from our own pinned, reproducible AI suite, 12 workloads: the Qwen3-4B to Llama-70B LLM ladder (llama.cpp, Q4_K_M), SDXL / Z-Image / FLUX-dev generation, FLUX-Kontext / Qwen-Edit editing, and LTX / Wan video, run first-party on rented hardware with under 0.5% run-to-run variance. Peak VRAM, power draw, temperature and tokens-per-watt are captured per workload. “Won’t fit” rows are real data: where a model exceeds the card’s VRAM at the tested precision we record a hard gate rather than silently dropping to a smaller quant. Measured 2026-07-11 · harness 2.0.0-standalone.

NVIDIA Quadro RTX 8000 specifications

ArchitectureTuring (TU102)
CUDA cores4,608
VRAM48GB GDDR6
Memory bus384-bit
Memory bandwidth672 GB/s
Boost clock1,770 MHz
TDP260 W
Process12nm
InterfacePCIe 3.0 x16
Release date2018-08-27
Launch MSRP$9,999

Verdict, capable, but 48GB sets the ceiling

NVIDIA Quadro RTX 8000 scores 8.2/100, #31 of 102. It ran 7 of 12; 1 exceeded its 48GB. Every figure here is our own measurement.

Relative performance: where the NVIDIA Quadro RTX 8000 lands

100% = this card, AI & Machine Learning headline metric (AI Score). #7 of 20 workstation cards in this vertical.

GPURelative%AI Score
NVIDIA RTX PRO 5000 Blackwell
385%31.6
NVIDIA RTX A6000
256%21
NVIDIA RTX PRO 4500 Blackwell
157%12.9
NVIDIA RTX A5500
102%8.4
NVIDIA Quadro RTX 8000
100%8.2
NVIDIA RTX PRO 4000 Blackwell
93%7.6
NVIDIA Quadro RTX 6000 (Turing)
90%7.4
NVIDIA RTX A5000
90%7.4
NVIDIA RTX 4500 Ada Generation
72%5.9

← All AI & Machine Learning GPU rankings

The silicon

Transistors18,600 million
Die size754 mm²
Process node12 nm
Fabricated byTSMC
Transistor density24.7 million per mm²

Denser than 69% of the 746 cards we have silicon data for. Density is the clearest measure of what a process node bought: a card that gained it without growing the die got its speed from the fab rather than the architecture.

Silicon figures from Wikipedia (CC BY-SA 4.0). Benchmarks on this page are our own. Compare every chip.

What this card can build

Whole-job timings, composed from our measured per-model results on this card.

WorkflowTimeEnergyBasis
Full codebase review21.4 min49.24 Whmeasured
24-frame storyboard54.9 min223 Whall 2 stages measured

Rent or buy?

This card is $9,999 to buy. The cheapest listed rate on Vast.ai is $0.255/hour, but that is the floor: we budget $0.306/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 32,676 GPU-hours. Below it you are paying for idle silicon.

How you would use itGPU-hours a yearRental cost a yearTime to break even
2 hours a day, hobby730$22344.8 years
8 hours a day, working on it2,920$89411.2 years
24/7, always-on agent8,760$2,6813.7 years

At hobby usage this card is very unlikely to pay for itself before it is superseded. Rent it. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.

Rental price

$0.255/hr+0.0% since 2026-09-30low $0.255 · high $0.255

Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.