NVIDIA Quadro RTX 6000 (Turing), AI & Machine Learning Benchmarks & Specs

24GB · AI Score 7.4/100 · first-party measured on 12 AI workloads

7.4 AI Score Includes estimates

Every number on this page is first-party: NVIDIA Quadro RTX 6000 (Turing) was run on our pinned 12-workload AI suite on 2026-07-11, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) NVIDIA Quadro RTX 6000 (Turing) delivers about 84.21 tokens/sec. Llama 3.3 70B does not fit. It needs roughly 42GB and this card has 24GB. For image generation, SDXL runs at 2.97 it/s, while FLUX.1-dev won't fit at BF16 (needs ~26GB). 5 of the 12 workloads won't fit on 24GB at the tested precision, Qwen3 32B, Llama 3.3 70B, FLUX.1-dev, FLUX.1 Kontext and others. We publish those as hard gates rather than quietly dropping to a smaller quant.

AI & Machine Learning benchmark results

Text Generation tok/s 5

Qwen3 4B125.23
Llama 3.1 8B84.21
Qwen2.5-Coder 14B46.54
WorkloadResultTelemetryData
Qwen3 4B125.23 tok/s
2.9 GB peak153 W49°C0.82 tok/WQ4_K_M
✓ Measured
Llama 3.1 8B84.21 tok/s
4.8 GB peak174 W57°C0.49 tok/WQ4_K_M
✓ Measured
Qwen2.5-Coder 14B46.54 tok/s
8.5 GB peak174 W66°C0.27 tok/WQ4_K_M
✓ Measured
Qwen3 32Bn/a tok/sestimatedEst.
Llama 3.3 70B✕ Won't fit needs ~46 GBVRAM-gated at this precision✓ Measured

Image Generation images/min 2

WorkloadResultTelemetryData
Stable Diffusion XL5.94 images/min
12.7 GB peak241 W76°C10.1 s/img
✓ Measured
FLUX.1 dev✕ Won't fit needs ~26 GBVRAM-gated at this precision✓ Measured

Image Editing images/min 2

WorkloadResultTelemetryData
FLUX.1 Kontext dev✕ Won't fit needs ~26 GBVRAM-gated at this precision✓ Measured
Qwen-Image-Edit✕ Won't fit needs ~42 GBVRAM-gated at this precision✓ Measured
How we measured this. Every result comes from our own pinned, reproducible AI suite, 12 workloads: the Qwen3-4B to Llama-70B LLM ladder (llama.cpp, Q4_K_M), SDXL / Z-Image / FLUX-dev generation, FLUX-Kontext / Qwen-Edit editing, and LTX / Wan video, run first-party on rented hardware with under 0.5% run-to-run variance. Peak VRAM, power draw, temperature and tokens-per-watt are captured per workload. “Won’t fit” rows are real data: where a model exceeds the card’s VRAM at the tested precision we record a hard gate rather than silently dropping to a smaller quant. Measured 2026-07-11 · harness 2.0.0-standalone.

NVIDIA Quadro RTX 6000 (Turing) specifications

ArchitectureTuring (TU102)
CUDA cores4,608
VRAM24GB GDDR6
Memory bus384-bit
Memory bandwidth672 GB/s
Boost clock1,770 MHz
TDP260 W
Process12nm
InterfacePCIe 3.0 x16
Release date2018-08-14
Launch MSRP$6,300

Verdict, capable, but 24GB sets the ceiling

NVIDIA Quadro RTX 6000 (Turing) scores 7.4/100, #36 of 102. It ran 4 of 12; 5 exceeded its 24GB. Every figure here is our own measurement.

Relative performance: where the NVIDIA Quadro RTX 6000 (Turing) lands

100% = this card, AI & Machine Learning headline metric (AI Score). #10 of 20 desktop cards in this vertical.

GPURelative%AI Score
AMD Radeon Pro W7900
173%12.8
NVIDIA Quadro RTX 8000
111%8.2
NVIDIA RTX A5500
103%7.6
NVIDIA RTX PRO 4000 Blackwell
103%7.6
NVIDIA Quadro RTX 6000 (Turing)
100%7.4
NVIDIA RTX A5000
100%7.4
AMD Radeon Pro W7800
84%6.2
NVIDIA RTX 4500 Ada Generation
80%5.9
AMD Radeon Pro W6800
74%5.5

← All AI & Machine Learning GPU rankings

What this card can build

Whole-job timings, composed from our measured per-model results on this card.

WorkflowTimeEnergyBasis
Full codebase review21.5 min62.2 Whmeasured

Can't run: 40-product photo shoot (needs FLUX.1 Kontext dev), 20 long-form articles (needs Llama 3.3 70B), Character sheet, 12 poses (needs FLUX.1 dev), 100-photo restoration batch (needs FLUX.1 Kontext dev), 100-photo restore and enlarge (needs FLUX.1 Kontext dev).