96GB · AI Score 44.9/100 · anchored estimate vs 51 measured cards
43.5 AI Score Includes estimates
We have not run NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition on our bench. These figures are anchored estimates, interpolated per workload against the 51 GPUs we did measure. On Llama 3.1 8B (Q4_K_M) NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition should deliver about 245.5 tokens/sec. Stepping up to Qwen3 32B it should hold roughly 67 tok/s. The full Llama 3.3 70B still runs, at about 33.1 tok/s. For image generation, SDXL should run near 10.9 it/s, and FLUX.1-dev at 2.52 it/s. All 12 workloads fit in 96GB. There is no model in our suite this card has to turn down.
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Qwen3-4B | 330.88 tok/s | 149 W50°CQ4_K_M | ✓ Measured |
| Llama-3.1-8B | 231.46 tok/s | 190 W51°CQ4_K_M | ✓ Measured |
| Qwen3 8B | 221.35 tok/s | 194 W52°CQ4_K_M | ✓ Measured |
| Gemma 4 12B | 135.12 tok/s | 204 W54°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder-14B | 125.82 tok/s | 208 W55°CQ4_K_M | ✓ Measured |
| Qwen3 14B | 129.38 tok/s | 216 W55°CQ4_K_M | ✓ Measured |
| gpt-oss-20b | 358.99 tok/s | 135 W53°CQ4_K_M | ✓ Measured |
| Qwen3 30B A3B | 311.38 tok/s | 137 W53°CQ4_K_M | ✓ Measured |
| Qwen3-32B | 59.97 tok/s | 237 W57°CQ4_K_M | ✓ Measured |
| Llama-3.3-70B | 29.38 tok/s | 248 W61°CQ4_K_M | ✓ Measured |
| GLM-4.5-Air | 122.9 tok/s | 155 W56°CQ4_K_M | ✓ Measured |
| Laguna-S-2.1 | 143.11 tok/s | 147 W56°CQ4_K_M | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Stable Diffusion XL | 21.8 images/min | estimated | Est. |
| FLUX.1 dev | 5.4 images/min | estimated | Est. |
| Z-Image Turbo | 12.15 images/min | estimated | Est. |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| FLUX.1 Kontext dev | 2.4 images/min | estimated | Est. |
| Qwen-Image-Edit | 2.06 images/min | estimated | Est. |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| LTX-Video (distilled) | 12.76 frames/s | estimated | Est. |
| Wan 2.2 5B (720p) | 0.81 frames/s | estimated | Est. |
| Architecture | Blackwell |
| CUDA cores | 24,064 |
| VRAM | 96GB GDDR7 ECC |
| Memory bus | 512-bit |
| Memory bandwidth | 1792 GB/s |
| Boost clock | 2,617 MHz |
| TDP | 300 W |
| Process | 4nm (TSMC 4N) |
| Interface | PCIe 5.0 x16 |
| Release date | 2025-03-18 |
| Launch MSRP | $8,565 |
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition scores 44.9/100, #12 of 102. It ran all 12 workloads. Figures are anchored estimates, not measurements, we flag that on every row.
100% = this card, AI & Machine Learning headline metric (AI Score). #2 of 20 workstation cards in this vertical.
| GPU | Relative | % | AI Score |
|---|---|---|---|
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 122% | 53.1 | |
| NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition | 100% | 43.5 | |
| NVIDIA RTX PRO 5000 Blackwell | 73% | 31.6 | |
| NVIDIA RTX A6000 | 48% | 21 | |
| NVIDIA RTX PRO 4500 Blackwell | 30% | 12.9 | |
| NVIDIA RTX A5500 | 19% | 8.4 |
← All AI & Machine Learning GPU rankings
| Transistors | 92,200 million |
| Die size | 750 mm² |
| Process node | 4 nm |
| Fabricated by | TSMC |
| Transistor density | 122.9 million per mm² |
Denser than 96% of the 746 cards we have silicon data for. Density is the clearest measure of what a process node bought: a card that gained it without growing the die got its speed from the fab rather than the architecture.
Silicon figures from Wikipedia (CC BY-SA 4.0). Benchmarks on this page are our own. Compare every chip.
Whole-job timings, composed from our measured per-model results on this card.
| Workflow | Time | Energy | Basis |
|---|---|---|---|
| 24-frame storyboard | 2.4 min | 1.54 Wh | estimate, 1 of 2 stages measured |
| 60-second AI short film | 2.9 min | 0.99 Wh | estimate, 1 of 3 stages measured |
| 6-panel comic page | 3.8 min | 0.77 Wh | estimate, 1 of 3 stages measured |
| Character sheet, 12 poses | 5.2 min | n/a | estimate, 0 of 2 stages measured |
| Full codebase review | 8 min | 27.59 Wh | measured |
| Short social clips | 11.1 min | 0.55 Wh | estimate, 1 of 3 stages measured |
| Long-form article batch | 15.9 min | 65.68 Wh | measured |
| Product photo shoot | 18.5 min | n/a | estimate, 0 of 2 stages measured |
| Photo restoration batch | 41.7 min | n/a | estimate, 0 of 1 stage measured |
This card is $8,565 to buy. The cheapest listed rate on Vast.ai is $1.069/hour, but that is the floor: we budget $1.283/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 6,677 GPU-hours. Below it you are paying for idle silicon.
| How you would use it | GPU-hours a year | Rental cost a year | Time to break even |
|---|---|---|---|
| 2 hours a day, hobby | 730 | $936 | 9.1 years |
| 8 hours a day, working on it | 2,920 | $3,746 | 2.3 years |
| 24/7, always-on agent | 8,760 | $11,237 | 9.1 months |
At hobby usage this card is very unlikely to pay for itself before it is superseded. Rent it. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.
Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.