96GB · AI Score 53.1/100 · first-party measured on 12 AI workloads
53.1 AI Score ✓ Measured
Every number on this page is first-party: NVIDIA RTX PRO 6000 Blackwell Workstation Edition was run on our pinned 12-workload AI suite on 2026-07-11, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) NVIDIA RTX PRO 6000 Blackwell Workstation Edition delivers about 258.37 tokens/sec. Stepping up to Qwen3 32B it holds roughly 70.13 tok/s. The full Llama 3.3 70B still runs, at about 34.87 tok/s. For image generation, SDXL runs at 13.97 it/s, and FLUX.1-dev at 3.23 it/s. All 12 workloads fit in 96GB. There is no model in our suite this card has to turn down. NVIDIA RTX PRO 6000 Blackwell Workstation Edition isn't a retail purchase for most people. It's rented by the hour. You can run this exact card on RunPod.
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| SmolLM2-135M | 1252.77 tok/s | 93 W46°CQ4_K_M | ✓ Measured |
| Qwen1.5-0.5B | 1025.58 tok/s | 103 W46°CQ4_K_M | ✓ Measured |
| LFM2.5-1.2B | 989.91 tok/s | 125 W53°CQ4_K_M | ✓ Measured |
| gemma-3-270m | 965 tok/s | 97 W43°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder-0.5B | 910.13 tok/s | 104 W45°CQ4_K_M | ✓ Measured |
| Qwen2.5-0.5B | 905.27 tok/s | 112 W43°CQ4_K_M | ✓ Measured |
| Llama 3.2 1B | 892.74 tok/s | 140 W38°CQ4_K_M | ✓ Measured |
| Qwen3-0.6B | 806.77 tok/s | 100 W44°CQ4_K_M | ✓ Measured |
| Qwen3 0.6B | 788.97 tok/s | 62 W36°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder-1.5B | 650.4 tok/s | 113 W54°CQ4_K_M | ✓ Measured |
| Qwen2-1.5B | 648.49 tok/s | 117 W48°CQ4_K_M | ✓ Measured |
| Qwen2.5-1.5B | 636.39 tok/s | 115 W44°CQ4_K_M | ✓ Measured |
| DeepSeek-R1 Distill 1.5B | 635.1 tok/s | 169 W45°CQ4_K_M | ✓ Measured |
| LFM2.5-8B-A1B | 617.09 tok/s | 129 W48°CQ4_K_M | ✓ Measured |
| Qwen3-1.7B | 593.5 tok/s | 143 W52°CQ4_K_M | ✓ Measured |
| Qwen3 1.7B | 575.77 tok/s | 105 W45°CQ4_K_M | ✓ Measured |
| gemma-3-1b | 567.1 tok/s | 128 W44°CQ4_K_M | ✓ Measured |
| Llama-3.2-3B-Instruct-uncensored | 440.15 tok/s | 145 W54°CQ4_K_M | ✓ Measured |
| Hermes-3-Llama-3.2-3B | 439.65 tok/s | 165 W49°CQ4_K_M | ✓ Measured |
| Llama 3.2 3B | 435.63 tok/s | 167 W47°CQ4_K_M | ✓ Measured |
| gemma-2-2b | 417.07 tok/s | 138 W47°CQ4_K_M | ✓ Measured |
| gemma-2-2b-it-abliterated | 416.68 tok/s | 143 W52°CQ4_K_M | ✓ Measured |
| phi-2 | 410.12 tok/s | 164 W49°CQ4_K_M | ✓ Measured |
| SmolLM3-3B | 409.73 tok/s | 140 W52°CQ4_K_M | ✓ Measured |
| SmolLM3 3B | 406.45 tok/s | 176 W46°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder-3B | 401.24 tok/s | 155 W55°CQ4_K_M | ✓ Measured |
| AI21-Jamba-Reasoning-3B | 400.02 tok/s | 161 W48°CQ4_K_M | ✓ Measured |
| Qwen2.5-3B | 398.72 tok/s | 153 W45°CQ4_K_M | ✓ Measured |
| Phi-4-mini | 393.81 tok/s | 194 W56°CQ4_K_M | ✓ Measured |
| Phi-4 Mini 3.8B | 391.04 tok/s | 172 W48°CQ4_K_M | ✓ Measured |
| gpt-oss-20b | 375.09 tok/s | 153 W44°CQ4_K_M | ✓ Measured |
| Qwen3 4B | 362.81 tok/s | 3.3 GB peak93 W39°C3.9 tok/WQ4_K_M | ✓ Measured |
| Phi-3.5-mini | 362.79 tok/s | 180 W46°CQ4_K_M | ✓ Measured |
| DeepSeek-Coder-V2-Lite | 352.52 tok/s | 155 W52°CQ4_K_M | ✓ Measured |
| Qwen3-4B-Instruct-2507 | 334.71 tok/s | 179 W51°CQ4_K_M | ✓ Measured |
| Qwen3-4B-Thinking-2507 | 334.65 tok/s | 192 W49°CQ4_K_M | ✓ Measured |
| Qwen3-4B-Instruct-2507 | 333.54 tok/s | 193 W52°CQ4_K_M | ✓ Measured |
| Nemotron-3-Nano-30B-A3B | 325.43 tok/s | 138 W50°CQ4_K_M | ✓ Measured |
| Qwen3-Coder 30B A3B | 317.68 tok/s | 111 W45°CQ4_K_M | ✓ Measured |
| Qwen3-30B-A3B | 311.76 tok/s | 145 W51°CQ4_K_M | ✓ Measured |
| Qwen3 30B A3B (Q3_K_M) | 309.61 tok/s | 145 W52°CQ3_K_M | ✓ Measured |
| Qwen3 30B A3B | 308.74 tok/s | 132 W47°CQ4_K_M | ✓ Measured |
| Gemma 3 4B | 297.73 tok/s | 169 W46°CQ4_K_M | ✓ Measured |
| Dolphin X1 Trinity Nano 6B | 282.02 tok/s | 124 W39°CQ4_K_M | ✓ Measured |
| Llama-2-7B | 263.84 tok/s | 241 W56°CQ4_K_M | ✓ Measured |
| Llama 3.1 8B | 258.37 tok/s | 5.2 GB peak241 W43°C1.07 tok/WQ4_K_M | ✓ Measured |
| Qwen2.5-Coder-7B-Instruct-abliterated | 257.07 tok/s | 204 W50°CQ4_K_M | ✓ Measured |
| Qwen2.5-7B | 256.63 tok/s | 228 W47°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder 7B | 256.11 tok/s | 195 W49°CQ4_K_M | ✓ Measured |
| DeepSeek-R1 Distill 7B | 256.04 tok/s | 169 W50°CQ4_K_M | ✓ Measured |
| KAT-Coder-V2.5-Dev | 253.35 tok/s | 152 W45°CQ4_K_M | ✓ Measured |
| Mistral-7B-Instruct-v0.2 | 253.34 tok/s | 244 W56°CQ4_K_M | ✓ Measured |
| Mistral-7B-Instruct-v0.3 | 253.2 tok/s | 222 W53°CQ4_K_M | ✓ Measured |
| Mistral-7B-Instruct-v0.1 | 252.55 tok/s | 238 W55°CQ4_K_M | ✓ Measured |
| Mistral 7B v0.3 | 247.33 tok/s | 175 W43°CQ4_K_M | ✓ Measured |
| L3-8B-Stheno-v3.2 | 239.84 tok/s | 224 W52°CQ4_K_M | ✓ Measured |
| DarkIdol-Llama-3.1-8B-Instruct-1.2-Uncensored | 239.45 tok/s | 227 W55°CQ4_K_M | ✓ Measured |
| Dolphin X1 8B | 239.4 tok/s | 202 W51°CQ4_K_M | ✓ Measured |
| dolphin-2.9-llama3-8b | 239.29 tok/s | 217 W53°CQ4_K_M | ✓ Measured |
| Dolphin 3.0 Llama 3.1 8B | 239.28 tok/s | 179 W50°CQ4_K_M | ✓ Measured |
| Ornith-1.0-35B | 237.7 tok/s | 155 W45°CQ4_K_M | ✓ Measured |
| Meta-Llama-3.1-8B | 237.69 tok/s | 213 W46°CQ4_K_M | ✓ Measured |
| DeepSeek-R1 Distill Llama 8B | 235.07 tok/s | 154 W44°CQ4_K_M | ✓ Measured |
| Qwen-AgentWorld-35B-A3B | 234.82 tok/s | 162 W45°CQ4_K_M | ✓ Measured |
| Josiefied-Qwen3-8B-abliterated-v1 | 227.15 tok/s | 213 W51°CQ4_K_M | ✓ Measured |
| Qwen3-8B | 226.88 tok/s | 223 W57°CQ4_K_M | ✓ Measured |
| DeepSeek-R1-0528-Qwen3-8B | 226.85 tok/s | 228 W48°CQ4_K_M | ✓ Measured |
| Qwen3 8B | 225.95 tok/s | 198 W49°CQ4_K_M | ✓ Measured |
| GLM-4.7-Flash | 215.61 tok/s | 163 W46°CQ4_K_M | ✓ Measured |
| Qwen3-Next-80B-A3B-Thinking | 214.04 tok/s | 135 W49°CQ4_K_M | ✓ Measured |
| Qwen3-Coder-Next | 211.5 tok/s | 147 W46°CQ4_K_M | ✓ Measured |
| Qwen3-Coder-Next-abliterated | 210.47 tok/s | 142 W47°CQ4_K_M | ✓ Measured |
| Qwen3-Next-80B-A3B | 210.17 tok/s | 136 W48°CQ4_K_M | ✓ Measured |
| Qwen3-Coder-Next | 205.69 tok/s | 145 W49°CQ4_K_M | ✓ Measured |
| Qwen3-Next-80B-A3B-Thinking | 203.29 tok/s | 144 W47°CQ4_K_M | ✓ Measured |
| Ornith-1.0-9B | 201.69 tok/s | 232 W49°CQ4_K_M | ✓ Measured |
| GLM-4.7-Flash-REAP-23B-A3B | 199.14 tok/s | 156 W50°CQ4_K_M | ✓ Measured |
| Phi-4 14B (Q3_K_M) | 165.43 tok/s | 286 W55°CQ3_K_M | ✓ Measured |
| Mistral-Nemo-Instruct-2407 | 164.28 tok/s | 262 W52°CQ4_K_M | ✓ Measured |
| NemoMix-Unleashed-12B | 164.21 tok/s | 252 W51°CQ4_K_M | ✓ Measured |
| gemma-2-9b | 156.29 tok/s | 240 W55°CQ4_K_M | ✓ Measured |
| Gemma 3 12B (Q3_K_M) | 153.93 tok/s | 257 W53°CQ3_K_M | ✓ Measured |
| DeepSeek-R1 Distill 14B (Q3_K_M) | 149.81 tok/s | 272 W54°CQ3_K_M | ✓ Measured |
| Qwen2.5-Coder 14B | 147.74 tok/s | 8.7 GB peak239 W47°C0.62 tok/WQ4_K_M | ✓ Measured |
| Phi-4 14B | 142.68 tok/s | 220 W49°CQ4_K_M | ✓ Measured |
| Hermes-4-14B | 139.56 tok/s | 277 W53°CQ4_K_M | ✓ Measured |
| Qwen3-14B | 139.45 tok/s | 263 W52°CQ4_K_M | ✓ Measured |
| Qwen3 14B | 139.01 tok/s | 175 W49°CQ4_K_M | ✓ Measured |
| Gemma 3 12B | 138.09 tok/s | 189 W45°CQ4_K_M | ✓ Measured |
| Gemma 4 12B | 137.73 tok/s | 201 W46°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder-14B-Instruct-abliterated | 135.88 tok/s | 281 W52°CQ4_K_M | ✓ Measured |
| EVA-Qwen2.5-14B-v0.2 | 135.83 tok/s | 271 W55°CQ4_K_M | ✓ Measured |
| Uncensored | 135.68 tok/s | 273 W53°CQ4_K_M | ✓ Measured |
| Qwen2.5-14B | 135.66 tok/s | 275 W49°CQ4_K_M | ✓ Measured |
| DeepSeek-R1 Distill 14B | 135.41 tok/s | 212 W49°CQ4_K_M | ✓ Measured |
| StarCoder2 15B | 122.09 tok/s | 211 W52°CQ4_K_M | ✓ Measured |
| Mistral Small 24B (Q3_K_M) | 108.88 tok/s | 334 W57°CQ3_K_M | ✓ Measured |
| Codestral 22B (Q3_K_M) | 107.67 tok/s | 338 W58°CQ3_K_M | ✓ Measured |
| Cydonia-24B-v4.3 | 93.76 tok/s | 301 W54°CQ4_K_M | ✓ Measured |
| Dolphin-Mistral-24B-Venice-Edition | 93.71 tok/s | 295 W53°CQ4_K_M | ✓ Measured |
| Dolphin Mistral 24B Venice | 93.67 tok/s | 157 W53°CQ4_K_M | ✓ Measured |
| Dolphin 3.0 R1 Mistral 24B | 93.62 tok/s | 199 W53°CQ4_K_M | ✓ Measured |
| Mistral Small 24B | 93.55 tok/s | 230 W50°CQ4_K_M | ✓ Measured |
| Devstral Small 24B | 93.52 tok/s | 136 W48°CQ4_K_M | ✓ Measured |
| Codestral 22B | 92.7 tok/s | 150 W51°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder 32B (Q3_K_M) | 75.78 tok/s | 351 W58°CQ3_K_M | ✓ Measured |
| Gemma 3 27B | 72.1 tok/s | 226 W60°CQ4_K_M | ✓ Measured |
| Qwen3 32B | 70.13 tok/s | 19.1 GB peak200 W51°C0.35 tok/WQ4_K_M | ✓ Measured |
| Qwen2.5-Coder 32B | 66.07 tok/s | 239 W57°CQ4_K_M | ✓ Measured |
| QwQ 32B | 66.07 tok/s | 218 W59°CQ4_K_M | ✓ Measured |
| DeepSeek-R1 Distill 32B | 66.06 tok/s | 260 W61°CQ4_K_M | ✓ Measured |
| DeepSeek-R1-Distill-Qwen-32B-abliterated | 66.03 tok/s | 320 W54°CQ4_K_M | ✓ Measured |
| Qwen2.5-32B | 66 tok/s | 321 W53°CQ4_K_M | ✓ Measured |
| Olmo-3.1-32B-Think | 65.71 tok/s | 321 W58°CQ4_K_M | ✓ Measured |
| Dolphin 2.9.1 Yi 1.5 34B | 63.74 tok/s | 186 W55°CQ4_K_M | ✓ Measured |
| Llama 3.3 70B | 34.87 tok/s | 40.1 GB peak193 W56°C0.18 tok/WQ4_K_M | ✓ Measured |
| Llama-3.3-70B-Instruct-abliterated | 32.53 tok/s | 363 W59°CQ4_K_M | ✓ Measured |
| DeepSeek-R1-Distill-Llama-70B | 32.52 tok/s | 358 W59°CQ4_K_M | ✓ Measured |
| Meta-Llama-3.1-70B | 32.52 tok/s | 352 W58°CQ4_K_M | ✓ Measured |
| Hermes-4-70B | 32.51 tok/s | 353 W57°CQ4_K_M | ✓ Measured |
| Qwen2.5-72B | 29.65 tok/s | 356 W57°CQ4_K_M | ✓ Measured |
| Laguna-XS-2.1 | ✕ Won't fit needs ~26 GB | VRAM-gated at this precision | ✓ Measured |
| Nanbeige4.2-3B | ✕ Won't fit needs ~4 GB | VRAM-gated at this precision | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| FLUX.1 Schnell | 42.41 images/min | 503 W53°C | ✓ Measured |
| Z-Image Turbo (1024px) | 31.43 images/min | 541 W62°C | ✓ Measured |
| Stable Diffusion XL | 27.94 images/min | 14.8 GB peak575 W55°C2.2 s/img | ✓ Measured |
| Z-Image Turbo | 15.6 images/min | 25.9 GB peak584 W57°C3.8 s/img | ✓ Measured |
| FLUX.1 dev | 6.92 images/min | 36.8 GB peak597 W66°C8.7 s/img | ✓ Measured |
| Krea 2 Turbo | 4.71 images/min | 491 W66°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| TinyLlama 1.1B LoRA | 18556.4 train tok/s | 271 W44°C | ✓ Measured |
| SmolLM2 1.7B LoRA | 16255.4 train tok/s | 330 W46°C | ✓ Measured |
| Qwen2.5 1.5B LoRA | 15753.9 train tok/s | 332 W44°C | ✓ Measured |
| Qwen2.5 7B LoRA | 6041.8 train tok/s | 513 W52°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| TinyLlama 1.1B served | 7678.7 serve tok/s | 175 W42°C | ✓ Measured |
| Qwen2.5 1.5B served | 5786.7 serve tok/s | 233 W43°C | ✓ Measured |
| SmolLM2 1.7B served | 5642 serve tok/s | 236 W42°C | ✓ Measured |
| Qwen2.5 7B served | 2282.8 serve tok/s | 362 W46°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| FLUX.1 Kontext dev | 3.09 images/min | 37.8 GB peak598 W76°C19.5 s/img | ✓ Measured |
| Qwen-Image-Edit | 2.64 images/min | 60.3 GB peak598 W74°C22.7 s/img | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Depth Anything V2 Small | 1374.58 images/min | 88 W34°C | ✓ Measured |
| Depth Anything V2 Large | 1328.11 images/min | 92 W35°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| SAM ViT-Base | 1079.89 images/min | 104 W38°C | ✓ Measured |
| SAM ViT-Huge | 193.47 images/min | 298 W47°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| LTX-Video (distilled) | 16.36 frames/s | 60.4 GB peak542 W63°C5.9 s/clip | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| BiRefNet | 1445.92 images/min | 100 W33°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Swin2SR 4x Upscaler | 52.83 images/min | 370 W44°C | ✓ Measured |
| Architecture | Blackwell |
| CUDA cores | 24,064 |
| VRAM | 96GB GDDR7 ECC |
| Memory bus | 512-bit |
| Memory bandwidth | 1792 GB/s |
| Boost clock | 2,617 MHz |
| TDP | 600 W |
| Process | 4nm (TSMC 4N) |
| Interface | PCIe 5.0 x16 |
| Release date | 2025-03-18 |
| Launch MSRP | $8,565 |
NVIDIA RTX PRO 6000 Blackwell Workstation Edition scores 53.1/100, #9 of 102. It ran all 12 workloads. Every figure here is our own measurement.
100% = this card, AI & Machine Learning headline metric (AI Score). #1 of 20 desktop cards in this vertical.
| GPU | Relative | % | AI Score |
|---|---|---|---|
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 100% | 53.1 | |
| NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition | 85% | 45.3 | |
| NVIDIA RTX PRO 5000 Blackwell | 60% | 31.6 | |
| NVIDIA RTX A6000 | 40% | 21 | |
| NVIDIA RTX PRO 4500 Blackwell | 24% | 12.9 |
← All AI & Machine Learning GPU rankings
Whole-job timings, composed from our measured per-model results on this card.
| Workflow | Time | Energy | Basis |
|---|---|---|---|
| 50-image depth pass | 4 s | 0.06 Wh | measured |
| 500-image masking run | 2.6 min | 12.84 Wh | measured |
| 24-frame storyboard | 3.6 min | 16.08 Wh | all 2 stages measured |
| 200-product catalogue cutout | 4 min | 23.56 Wh | all 2 stages measured |
| 60-second AI short film | 4.5 min | 19.26 Wh | all 3 stages measured |
| 6-panel comic page | 5.1 min | 28.55 Wh | all 3 stages measured |
| Character sheet, 12 poses | 6.1 min | 40.17 Wh | all 2 stages measured |
| Full codebase review | 6.8 min | 27 Wh | measured |
| 20 long-form articles | 13.4 min | 42.96 Wh | measured |
| 40-product photo shoot | 16.7 min | 142.82 Wh | all 2 stages measured |
| 40-product shoot, start to finish | 17.6 min | 147.54 Wh | all 4 stages measured |
| 100-photo restoration batch | 33.4 min | 322.75 Wh | measured |
| 100-photo restore and enlarge | 35.3 min | 334.41 Wh | all 2 stages measured |