NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition, AI & Machine Learning Benchmarks & Specs

96GB · AI Score 44.9/100 · anchored estimate vs 51 measured cards

43.5 AI Score Includes estimates

We have not run NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition on our bench. These figures are anchored estimates, interpolated per workload against the 51 GPUs we did measure. On Llama 3.1 8B (Q4_K_M) NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition should deliver about 245.5 tokens/sec. Stepping up to Qwen3 32B it should hold roughly 67 tok/s. The full Llama 3.3 70B still runs, at about 33.1 tok/s. For image generation, SDXL should run near 10.9 it/s, and FLUX.1-dev at 2.52 it/s. All 12 workloads fit in 96GB. There is no model in our suite this card has to turn down.

AI & Machine Learning benchmark results

Text Generation tok/s 12

gpt-oss-20b358.99
Qwen3-4B330.88
Qwen3 30B A3B311.38
Llama-3.1-8B231.46
Qwen3 8B221.35
Laguna-S-2.1143.11
Gemma 4 12B135.12
Qwen3 14B129.38
Qwen2.5-Coder-14B125.82
GLM-4.5-Air122.9
Qwen3-32B59.97
Llama-3.3-70B29.38
WorkloadResultTelemetryData
Qwen3-4B330.88 tok/s
149 W50°CQ4_K_M
✓ Measured
Llama-3.1-8B231.46 tok/s
190 W51°CQ4_K_M
✓ Measured
Qwen3 8B221.35 tok/s
194 W52°CQ4_K_M
✓ Measured
Gemma 4 12B135.12 tok/s
204 W54°CQ4_K_M
✓ Measured
Qwen2.5-Coder-14B125.82 tok/s
208 W55°CQ4_K_M
✓ Measured
Qwen3 14B129.38 tok/s
216 W55°CQ4_K_M
✓ Measured
gpt-oss-20b358.99 tok/s
135 W53°CQ4_K_M
✓ Measured
Qwen3 30B A3B311.38 tok/s
137 W53°CQ4_K_M
✓ Measured
Qwen3-32B59.97 tok/s
237 W57°CQ4_K_M
✓ Measured
Llama-3.3-70B29.38 tok/s
248 W61°CQ4_K_M
✓ Measured
GLM-4.5-Air122.9 tok/s
155 W56°CQ4_K_M
✓ Measured
Laguna-S-2.1143.11 tok/s
147 W56°CQ4_K_M
✓ Measured

Image Generation images/min 3

Stable Diffusion XL21.8
Z-Image Turbo12.15
FLUX.1 dev5.4
WorkloadResultTelemetryData
Stable Diffusion XL21.8 images/minestimatedEst.
FLUX.1 dev5.4 images/minestimatedEst.
Z-Image Turbo12.15 images/minestimatedEst.

Image Editing images/min 2

WorkloadResultTelemetryData
FLUX.1 Kontext dev2.4 images/minestimatedEst.
Qwen-Image-Edit2.06 images/minestimatedEst.

Video Generation frames/s 2

WorkloadResultTelemetryData
LTX-Video (distilled)12.76 frames/sestimatedEst.
Wan 2.2 5B (720p)0.81 frames/sestimatedEst.
How this estimate is derived. This card hasn’t been through our bench yet, so its numbers are anchored estimates, interpolated from the 51 first-party measured cards (Sibling-anchored: same GB202 die, power-capped (~300W); LLM ×0.95 (memory-bound), diffusion/video ×0.78 (power-bound)). The VRAM “won’t fit” gates are exact, since they’re pure capacity limits. Estimates are replaced with measured data as more silicon goes through the bench. Full methodology →

NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition specifications

ArchitectureBlackwell
CUDA cores24,064
VRAM96GB GDDR7 ECC
Memory bus512-bit
Memory bandwidth1792 GB/s
Boost clock2,617 MHz
TDP300 W
Process4nm (TSMC 4N)
InterfacePCIe 5.0 x16
Release date2025-03-18
Launch MSRP$8,565

Verdict, nothing in our suite slows it down

NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition scores 44.9/100, #12 of 102. It ran all 12 workloads. Figures are anchored estimates, not measurements, we flag that on every row.

Relative performance: where the NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition lands

100% = this card, AI & Machine Learning headline metric (AI Score). #2 of 20 workstation cards in this vertical.

GPURelative%AI Score
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
122%53.1
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
100%43.5
NVIDIA RTX PRO 5000 Blackwell
73%31.6
NVIDIA RTX A6000
48%21
NVIDIA RTX PRO 4500 Blackwell
30%12.9
NVIDIA RTX A5500
19%8.4

← All AI & Machine Learning GPU rankings

The silicon

Transistors92,200 million
Die size750 mm²
Process node4 nm
Fabricated byTSMC
Transistor density122.9 million per mm²

Denser than 96% of the 746 cards we have silicon data for. Density is the clearest measure of what a process node bought: a card that gained it without growing the die got its speed from the fab rather than the architecture.

Silicon figures from Wikipedia (CC BY-SA 4.0). Benchmarks on this page are our own. Compare every chip.

What this card can build

Whole-job timings, composed from our measured per-model results on this card.

WorkflowTimeEnergyBasis
24-frame storyboard2.4 min1.54 Whestimate, 1 of 2 stages measured
60-second AI short film2.9 min0.99 Whestimate, 1 of 3 stages measured
6-panel comic page3.8 min0.77 Whestimate, 1 of 3 stages measured
Character sheet, 12 poses5.2 minn/aestimate, 0 of 2 stages measured
Full codebase review8 min27.59 Whmeasured
Short social clips11.1 min0.55 Whestimate, 1 of 3 stages measured
Long-form article batch15.9 min65.68 Whmeasured
Product photo shoot18.5 minn/aestimate, 0 of 2 stages measured
Photo restoration batch41.7 minn/aestimate, 0 of 1 stage measured

Rent or buy?

This card is $8,565 to buy. The cheapest listed rate on Vast.ai is $1.069/hour, but that is the floor: we budget $1.283/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 6,677 GPU-hours. Below it you are paying for idle silicon.

How you would use itGPU-hours a yearRental cost a yearTime to break even
2 hours a day, hobby730$9369.1 years
8 hours a day, working on it2,920$3,7462.3 years
24/7, always-on agent8,760$11,2379.1 months

At hobby usage this card is very unlikely to pay for itself before it is superseded. Rent it. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.

Rental price

$1.069/hr-6.0% since 2026-10-01low $1.001 · high $1.137

Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.