NVIDIA RTX PRO 6000 Blackwell Server Edition, AI & Machine Learning Benchmarks & Specs

96GB · AI Score 49.9/100 · first-party measured on 12 AI workloads

49.9 AI Score ✓ Measured

Every number on this page is first-party: NVIDIA RTX PRO 6000 Blackwell Server Edition was run on our pinned 12-workload AI suite on 2026-07-10, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) NVIDIA RTX PRO 6000 Blackwell Server Edition delivers about 233.8 tokens/sec. Stepping up to Qwen3 32B it holds roughly 63.9 tok/s. The full Llama 3.3 70B still runs, at about 32.08 tok/s. For image generation, SDXL runs at 13.16 it/s, and FLUX.1-dev at 3.15 it/s. All 12 workloads fit in 96GB. There is no model in our suite this card has to turn down. NVIDIA RTX PRO 6000 Blackwell Server Edition isn't a retail purchase for most people. It's rented by the hour. You can run this exact card on RunPod.

AI & Machine Learning benchmark results

Text Generation tok/s 5

Qwen3 4B315.72
Llama 3.1 8B233.8
Qwen2.5-Coder 14B135.07
Qwen3 32B63.9
Llama 3.3 70B32.08
WorkloadResultTelemetryData
Qwen3 4B315.72 tok/s
3 GB peak172 W43°C1.83 tok/WQ4_K_M
✓ Measured
Llama 3.1 8B233.8 tok/s
5.2 GB peak217 W45°C1.08 tok/WQ4_K_M
✓ Measured
Qwen2.5-Coder 14B135.07 tok/s
9 GB peak261 W48°C0.52 tok/WQ4_K_M
✓ Measured
Qwen3 32B63.9 tok/s
19 GB peak218 W51°C0.29 tok/WQ4_K_M
✓ Measured
Llama 3.3 70B32.08 tok/s
40.1 GB peak251 W55°C0.13 tok/WQ4_K_M
✓ Measured

Image Generation images/min 3

Stable Diffusion XL26.32
Z-Image Turbo14.775
FLUX.1 dev6.75
WorkloadResultTelemetryData
Stable Diffusion XL26.32 images/min
14.7 GB peak509 W59°C2.3 s/img
✓ Measured
Z-Image Turbo14.78 images/min
25.9 GB peak531 W63°C4.1 s/img
✓ Measured
FLUX.1 dev6.75 images/min
36.8 GB peak544 W67°C8.9 s/img
✓ Measured

Image Editing images/min 2

WorkloadResultTelemetryData
FLUX.1 Kontext dev3.02 images/min
35.6 GB peak547 W72°C19.9 s/img
✓ Measured
Qwen-Image-Edit2.62 images/min
60.3 GB peak567 W73°C23 s/img
✓ Measured

Video Generation frames/s 2

WorkloadResultTelemetryData
Wan 2.2 5B (720p)0.88 frames/s
37.6 GB peak531 W72°C55.4 s/clip
✓ Measured
LTX-Video (distilled)15.64 frames/s
60.4 GB peak524 W67°C6.2 s/clip
✓ Measured
How we measured this. Every result comes from our own pinned, reproducible AI suite, 12 workloads: the Qwen3-4B to Llama-70B LLM ladder (llama.cpp, Q4_K_M), SDXL / Z-Image / FLUX-dev generation, FLUX-Kontext / Qwen-Edit editing, and LTX / Wan video, run first-party on rented hardware with under 0.5% run-to-run variance. Peak VRAM, power draw, temperature and tokens-per-watt are captured per workload. “Won’t fit” rows are real data: where a model exceeds the card’s VRAM at the tested precision we record a hard gate rather than silently dropping to a smaller quant. Measured 2026-07-10 · harness 2.0.0.

NVIDIA RTX PRO 6000 Blackwell Server Edition specifications

ArchitectureBlackwell
CUDA cores24,064
VRAM96GB GDDR7 ECC
Memory bus512-bit
Memory bandwidth1792 GB/s
Boost clock2,610 MHz
TDP600 W
ProcessTSMC 4N
InterfacePCIe 5.0 x16
Release date2025-03-18
Launch MSRP$8,565

Verdict, NVIDIA RTX PRO 6000 Blackwell Server Edition on real AI workloads

NVIDIA RTX PRO 6000 Blackwell Server Edition scores 49.9/100, #10 of 102. It ran all 12 workloads. Every figure here is our own measurement.

Relative performance: where the NVIDIA RTX PRO 6000 Blackwell Server Edition lands

100% = this card, AI & Machine Learning headline metric (AI Score). #9 of 21 datacenter cards in this vertical.

GPURelative%AI Score
NVIDIA H200
130%65
NVIDIA H100 80GB HBM3
125%62.6
NVIDIA H800 80GB
125%62.6
NVIDIA H100 NVL
116%58.1
NVIDIA RTX PRO 6000 Blackwell Server Edition
100%49.9
NVIDIA H100 PCIe
100%49.8
NVIDIA A100 80GB SXM4
66%33.1
NVIDIA A800 80GB
66%33.1
NVIDIA A100 80GB PCIe
64%31.7

← All AI & Machine Learning GPU rankings

The silicon

Transistors92,200 million
Die size750 mm²
Process node4 nm
Fabricated byTSMC
Transistor density122.9 million per mm²

Denser than 96% of the 746 cards we have silicon data for. Density is the clearest measure of what a process node bought: a card that gained it without growing the die got its speed from the fab rather than the architecture.

Silicon figures from Wikipedia (CC BY-SA 4.0). Benchmarks on this page are our own. Compare every chip.

What this card can build

Whole-job timings, composed from our measured per-model results on this card.

WorkflowTimeEnergyBasis
24-frame storyboard2.3 min15.7 Whall 2 stages measured
60-second AI short film2.7 min19.22 Whall 3 stages measured
6-panel comic page3.6 min26.83 Whall 3 stages measured
Character sheet, 12 poses4.6 min37.55 Whall 2 stages measured
Full codebase review7.4 min32.21 Whmeasured
Short social clips10.8 min88.63 Whall 3 stages measured
Long-form article batch14.5 min60.83 Whmeasured
Product photo shoot15.2 min133.59 Whall 2 stages measured
Photo restoration batch33.4 min301.72 Whmeasured

Rent or buy?

This card is $8,565 to buy. The cheapest listed rate on Vast.ai is $1.296/hour, but that is the floor: we budget $1.555/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 5,507 GPU-hours. Below it you are paying for idle silicon.

How you would use itGPU-hours a yearRental cost a yearTime to break even
2 hours a day, hobby730$1,1357.5 years
8 hours a day, working on it2,920$4,5411.9 years
24/7, always-on agent8,760$13,6247.5 months

At hobby usage this card is very unlikely to pay for itself before it is superseded. Rent it. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.

Rental price

$1.296/hr-3.0% since 2026-10-01low $1.004 · high $1.336

Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.