NVIDIA RTX PRO 6000 Blackwell Workstation Edition, AI & Machine Learning Benchmarks & Specs

96GB · AI Score 53.1/100 · first-party measured on 12 AI workloads

53.1 AI Score ✓ Measured

Every number on this page is first-party: NVIDIA RTX PRO 6000 Blackwell Workstation Edition was run on our pinned 12-workload AI suite on 2026-07-11, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) NVIDIA RTX PRO 6000 Blackwell Workstation Edition delivers about 258.37 tokens/sec. Stepping up to Qwen3 32B it holds roughly 70.13 tok/s. The full Llama 3.3 70B still runs, at about 34.87 tok/s. For image generation, SDXL runs at 13.97 it/s, and FLUX.1-dev at 3.23 it/s. All 12 workloads fit in 96GB. There is no model in our suite this card has to turn down. NVIDIA RTX PRO 6000 Blackwell Workstation Edition isn't a retail purchase for most people. It's rented by the hour. You can run this exact card on RunPod.

AI & Machine Learning benchmark results

Text Generation tok/s 123

SmolLM2-135M1252.77
Qwen1.5-0.5B1025.58
LFM2.5-1.2B989.91
gemma-3-270m965
Qwen2.5-Coder-0.5B910.13
Qwen2.5-0.5B905.27
Llama 3.2 1B892.74
Qwen3-0.6B806.77
Qwen3 0.6B788.97
Qwen2.5-Coder-1.5B650.4
Qwen2-1.5B648.49
Qwen2.5-1.5B636.39
WorkloadResultTelemetryData
SmolLM2-135M1252.77 tok/s
93 W46°CQ4_K_M
✓ Measured
Qwen1.5-0.5B1025.58 tok/s
103 W46°CQ4_K_M
✓ Measured
LFM2.5-1.2B989.91 tok/s
125 W53°CQ4_K_M
✓ Measured
gemma-3-270m965 tok/s
97 W43°CQ4_K_M
✓ Measured
Qwen2.5-Coder-0.5B910.13 tok/s
104 W45°CQ4_K_M
✓ Measured
Qwen2.5-0.5B905.27 tok/s
112 W43°CQ4_K_M
✓ Measured
Llama 3.2 1B892.74 tok/s
140 W38°CQ4_K_M
✓ Measured
Qwen3-0.6B806.77 tok/s
100 W44°CQ4_K_M
✓ Measured
Qwen3 0.6B788.97 tok/s
62 W36°CQ4_K_M
✓ Measured
Qwen2.5-Coder-1.5B650.4 tok/s
113 W54°CQ4_K_M
✓ Measured
Qwen2-1.5B648.49 tok/s
117 W48°CQ4_K_M
✓ Measured
Qwen2.5-1.5B636.39 tok/s
115 W44°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 1.5B635.1 tok/s
169 W45°CQ4_K_M
✓ Measured
LFM2.5-8B-A1B617.09 tok/s
129 W48°CQ4_K_M
✓ Measured
Qwen3-1.7B593.5 tok/s
143 W52°CQ4_K_M
✓ Measured
Qwen3 1.7B575.77 tok/s
105 W45°CQ4_K_M
✓ Measured
gemma-3-1b567.1 tok/s
128 W44°CQ4_K_M
✓ Measured
Llama-3.2-3B-Instruct-uncensored440.15 tok/s
145 W54°CQ4_K_M
✓ Measured
Hermes-3-Llama-3.2-3B439.65 tok/s
165 W49°CQ4_K_M
✓ Measured
Llama 3.2 3B435.63 tok/s
167 W47°CQ4_K_M
✓ Measured
gemma-2-2b417.07 tok/s
138 W47°CQ4_K_M
✓ Measured
gemma-2-2b-it-abliterated416.68 tok/s
143 W52°CQ4_K_M
✓ Measured
phi-2410.12 tok/s
164 W49°CQ4_K_M
✓ Measured
SmolLM3-3B409.73 tok/s
140 W52°CQ4_K_M
✓ Measured
SmolLM3 3B406.45 tok/s
176 W46°CQ4_K_M
✓ Measured
Qwen2.5-Coder-3B401.24 tok/s
155 W55°CQ4_K_M
✓ Measured
AI21-Jamba-Reasoning-3B400.02 tok/s
161 W48°CQ4_K_M
✓ Measured
Qwen2.5-3B398.72 tok/s
153 W45°CQ4_K_M
✓ Measured
Phi-4-mini393.81 tok/s
194 W56°CQ4_K_M
✓ Measured
Phi-4 Mini 3.8B391.04 tok/s
172 W48°CQ4_K_M
✓ Measured
gpt-oss-20b375.09 tok/s
153 W44°CQ4_K_M
✓ Measured
Qwen3 4B362.81 tok/s
3.3 GB peak93 W39°C3.9 tok/WQ4_K_M
✓ Measured
Phi-3.5-mini362.79 tok/s
180 W46°CQ4_K_M
✓ Measured
DeepSeek-Coder-V2-Lite352.52 tok/s
155 W52°CQ4_K_M
✓ Measured
Qwen3-4B-Instruct-2507334.71 tok/s
179 W51°CQ4_K_M
✓ Measured
Qwen3-4B-Thinking-2507334.65 tok/s
192 W49°CQ4_K_M
✓ Measured
Qwen3-4B-Instruct-2507333.54 tok/s
193 W52°CQ4_K_M
✓ Measured
Nemotron-3-Nano-30B-A3B325.43 tok/s
138 W50°CQ4_K_M
✓ Measured
Qwen3-Coder 30B A3B317.68 tok/s
111 W45°CQ4_K_M
✓ Measured
Qwen3-30B-A3B311.76 tok/s
145 W51°CQ4_K_M
✓ Measured
Qwen3 30B A3B (Q3_K_M)309.61 tok/s
145 W52°CQ3_K_M
✓ Measured
Qwen3 30B A3B308.74 tok/s
132 W47°CQ4_K_M
✓ Measured
Gemma 3 4B297.73 tok/s
169 W46°CQ4_K_M
✓ Measured
Dolphin X1 Trinity Nano 6B282.02 tok/s
124 W39°CQ4_K_M
✓ Measured
Llama-2-7B263.84 tok/s
241 W56°CQ4_K_M
✓ Measured
Llama 3.1 8B258.37 tok/s
5.2 GB peak241 W43°C1.07 tok/WQ4_K_M
✓ Measured
Qwen2.5-Coder-7B-Instruct-abliterated257.07 tok/s
204 W50°CQ4_K_M
✓ Measured
Qwen2.5-7B256.63 tok/s
228 W47°CQ4_K_M
✓ Measured
Qwen2.5-Coder 7B256.11 tok/s
195 W49°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 7B256.04 tok/s
169 W50°CQ4_K_M
✓ Measured
KAT-Coder-V2.5-Dev253.35 tok/s
152 W45°CQ4_K_M
✓ Measured
Mistral-7B-Instruct-v0.2253.34 tok/s
244 W56°CQ4_K_M
✓ Measured
Mistral-7B-Instruct-v0.3253.2 tok/s
222 W53°CQ4_K_M
✓ Measured
Mistral-7B-Instruct-v0.1252.55 tok/s
238 W55°CQ4_K_M
✓ Measured
Mistral 7B v0.3247.33 tok/s
175 W43°CQ4_K_M
✓ Measured
L3-8B-Stheno-v3.2239.84 tok/s
224 W52°CQ4_K_M
✓ Measured
DarkIdol-Llama-3.1-8B-Instruct-1.2-Uncensored239.45 tok/s
227 W55°CQ4_K_M
✓ Measured
Dolphin X1 8B239.4 tok/s
202 W51°CQ4_K_M
✓ Measured
dolphin-2.9-llama3-8b239.29 tok/s
217 W53°CQ4_K_M
✓ Measured
Dolphin 3.0 Llama 3.1 8B239.28 tok/s
179 W50°CQ4_K_M
✓ Measured
Ornith-1.0-35B237.7 tok/s
155 W45°CQ4_K_M
✓ Measured
Meta-Llama-3.1-8B237.69 tok/s
213 W46°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill Llama 8B235.07 tok/s
154 W44°CQ4_K_M
✓ Measured
Qwen-AgentWorld-35B-A3B234.82 tok/s
162 W45°CQ4_K_M
✓ Measured
Josiefied-Qwen3-8B-abliterated-v1227.15 tok/s
213 W51°CQ4_K_M
✓ Measured
Qwen3-8B226.88 tok/s
223 W57°CQ4_K_M
✓ Measured
DeepSeek-R1-0528-Qwen3-8B226.85 tok/s
228 W48°CQ4_K_M
✓ Measured
Qwen3 8B225.95 tok/s
198 W49°CQ4_K_M
✓ Measured
GLM-4.7-Flash215.61 tok/s
163 W46°CQ4_K_M
✓ Measured
Qwen3-Next-80B-A3B-Thinking214.04 tok/s
135 W49°CQ4_K_M
✓ Measured
Qwen3-Coder-Next211.5 tok/s
147 W46°CQ4_K_M
✓ Measured
Qwen3-Coder-Next-abliterated210.47 tok/s
142 W47°CQ4_K_M
✓ Measured
Qwen3-Next-80B-A3B210.17 tok/s
136 W48°CQ4_K_M
✓ Measured
Qwen3-Coder-Next205.69 tok/s
145 W49°CQ4_K_M
✓ Measured
Qwen3-Next-80B-A3B-Thinking203.29 tok/s
144 W47°CQ4_K_M
✓ Measured
Ornith-1.0-9B201.69 tok/s
232 W49°CQ4_K_M
✓ Measured
GLM-4.7-Flash-REAP-23B-A3B199.14 tok/s
156 W50°CQ4_K_M
✓ Measured
Phi-4 14B (Q3_K_M)165.43 tok/s
286 W55°CQ3_K_M
✓ Measured
Mistral-Nemo-Instruct-2407164.28 tok/s
262 W52°CQ4_K_M
✓ Measured
NemoMix-Unleashed-12B164.21 tok/s
252 W51°CQ4_K_M
✓ Measured
gemma-2-9b156.29 tok/s
240 W55°CQ4_K_M
✓ Measured
Gemma 3 12B (Q3_K_M)153.93 tok/s
257 W53°CQ3_K_M
✓ Measured
DeepSeek-R1 Distill 14B (Q3_K_M)149.81 tok/s
272 W54°CQ3_K_M
✓ Measured
Qwen2.5-Coder 14B147.74 tok/s
8.7 GB peak239 W47°C0.62 tok/WQ4_K_M
✓ Measured
Phi-4 14B142.68 tok/s
220 W49°CQ4_K_M
✓ Measured
Hermes-4-14B139.56 tok/s
277 W53°CQ4_K_M
✓ Measured
Qwen3-14B139.45 tok/s
263 W52°CQ4_K_M
✓ Measured
Qwen3 14B139.01 tok/s
175 W49°CQ4_K_M
✓ Measured
Gemma 3 12B138.09 tok/s
189 W45°CQ4_K_M
✓ Measured
Gemma 4 12B137.73 tok/s
201 W46°CQ4_K_M
✓ Measured
Qwen2.5-Coder-14B-Instruct-abliterated135.88 tok/s
281 W52°CQ4_K_M
✓ Measured
EVA-Qwen2.5-14B-v0.2135.83 tok/s
271 W55°CQ4_K_M
✓ Measured
Uncensored135.68 tok/s
273 W53°CQ4_K_M
✓ Measured
Qwen2.5-14B135.66 tok/s
275 W49°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 14B135.41 tok/s
212 W49°CQ4_K_M
✓ Measured
StarCoder2 15B122.09 tok/s
211 W52°CQ4_K_M
✓ Measured
Mistral Small 24B (Q3_K_M)108.88 tok/s
334 W57°CQ3_K_M
✓ Measured
Codestral 22B (Q3_K_M)107.67 tok/s
338 W58°CQ3_K_M
✓ Measured
Cydonia-24B-v4.393.76 tok/s
301 W54°CQ4_K_M
✓ Measured
Dolphin-Mistral-24B-Venice-Edition93.71 tok/s
295 W53°CQ4_K_M
✓ Measured
Dolphin Mistral 24B Venice93.67 tok/s
157 W53°CQ4_K_M
✓ Measured
Dolphin 3.0 R1 Mistral 24B93.62 tok/s
199 W53°CQ4_K_M
✓ Measured
Mistral Small 24B93.55 tok/s
230 W50°CQ4_K_M
✓ Measured
Devstral Small 24B93.52 tok/s
136 W48°CQ4_K_M
✓ Measured
Codestral 22B92.7 tok/s
150 W51°CQ4_K_M
✓ Measured
Qwen2.5-Coder 32B (Q3_K_M)75.78 tok/s
351 W58°CQ3_K_M
✓ Measured
Gemma 3 27B72.1 tok/s
226 W60°CQ4_K_M
✓ Measured
Qwen3 32B70.13 tok/s
19.1 GB peak200 W51°C0.35 tok/WQ4_K_M
✓ Measured
Qwen2.5-Coder 32B66.07 tok/s
239 W57°CQ4_K_M
✓ Measured
QwQ 32B66.07 tok/s
218 W59°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 32B66.06 tok/s
260 W61°CQ4_K_M
✓ Measured
DeepSeek-R1-Distill-Qwen-32B-abliterated66.03 tok/s
320 W54°CQ4_K_M
✓ Measured
Qwen2.5-32B66 tok/s
321 W53°CQ4_K_M
✓ Measured
Olmo-3.1-32B-Think65.71 tok/s
321 W58°CQ4_K_M
✓ Measured
Dolphin 2.9.1 Yi 1.5 34B63.74 tok/s
186 W55°CQ4_K_M
✓ Measured
Llama 3.3 70B34.87 tok/s
40.1 GB peak193 W56°C0.18 tok/WQ4_K_M
✓ Measured
Llama-3.3-70B-Instruct-abliterated32.53 tok/s
363 W59°CQ4_K_M
✓ Measured
DeepSeek-R1-Distill-Llama-70B32.52 tok/s
358 W59°CQ4_K_M
✓ Measured
Meta-Llama-3.1-70B32.52 tok/s
352 W58°CQ4_K_M
✓ Measured
Hermes-4-70B32.51 tok/s
353 W57°CQ4_K_M
✓ Measured
Qwen2.5-72B29.65 tok/s
356 W57°CQ4_K_M
✓ Measured
Laguna-XS-2.1✕ Won't fit needs ~26 GBVRAM-gated at this precision✓ Measured
Nanbeige4.2-3B✕ Won't fit needs ~4 GBVRAM-gated at this precision✓ Measured

Image Generation images/min 6

FLUX.1 Schnell42.41
Z-Image Turbo (1024px)31.43
Stable Diffusion XL27.94
Z-Image Turbo15.6
FLUX.1 dev6.921
Krea 2 Turbo4.71
WorkloadResultTelemetryData
FLUX.1 Schnell42.41 images/min
503 W53°C
✓ Measured
Z-Image Turbo (1024px)31.43 images/min
541 W62°C
✓ Measured
Stable Diffusion XL27.94 images/min
14.8 GB peak575 W55°C2.2 s/img
✓ Measured
Z-Image Turbo15.6 images/min
25.9 GB peak584 W57°C3.8 s/img
✓ Measured
FLUX.1 dev6.92 images/min
36.8 GB peak597 W66°C8.7 s/img
✓ Measured
Krea 2 Turbo4.71 images/min
491 W66°C
✓ Measured

Fine-Tuning train tok/s 4

TinyLlama 1.1B LoRA18556.4
SmolLM2 1.7B LoRA16255.4
Qwen2.5 1.5B LoRA15753.9
Qwen2.5 7B LoRA6041.8
WorkloadResultTelemetryData
TinyLlama 1.1B LoRA18556.4 train tok/s
271 W44°C
✓ Measured
SmolLM2 1.7B LoRA16255.4 train tok/s
330 W46°C
✓ Measured
Qwen2.5 1.5B LoRA15753.9 train tok/s
332 W44°C
✓ Measured
Qwen2.5 7B LoRA6041.8 train tok/s
513 W52°C
✓ Measured

LLM Serving serve tok/s 4

TinyLlama 1.1B served7678.7
Qwen2.5 1.5B served5786.7
SmolLM2 1.7B served5642
Qwen2.5 7B served2282.8
WorkloadResultTelemetryData
TinyLlama 1.1B served7678.7 serve tok/s
175 W42°C
✓ Measured
Qwen2.5 1.5B served5786.7 serve tok/s
233 W43°C
✓ Measured
SmolLM2 1.7B served5642 serve tok/s
236 W42°C
✓ Measured
Qwen2.5 7B served2282.8 serve tok/s
362 W46°C
✓ Measured

Image Editing images/min 2

WorkloadResultTelemetryData
FLUX.1 Kontext dev3.09 images/min
37.8 GB peak598 W76°C19.5 s/img
✓ Measured
Qwen-Image-Edit2.64 images/min
60.3 GB peak598 W74°C22.7 s/img
✓ Measured

Depth Estimation images/min 2

WorkloadResultTelemetryData
Depth Anything V2 Small1374.58 images/min
88 W34°C
✓ Measured
Depth Anything V2 Large1328.11 images/min
92 W35°C
✓ Measured

Segmentation images/min 2

WorkloadResultTelemetryData
SAM ViT-Base1079.89 images/min
104 W38°C
✓ Measured
SAM ViT-Huge193.47 images/min
298 W47°C
✓ Measured

Video Generation frames/s 1

WorkloadResultTelemetryData
LTX-Video (distilled)16.36 frames/s
60.4 GB peak542 W63°C5.9 s/clip
✓ Measured

Background Removal images/min 1

WorkloadResultTelemetryData
BiRefNet1445.92 images/min
100 W33°C
✓ Measured

Upscaling images/min 1

WorkloadResultTelemetryData
Swin2SR 4x Upscaler52.83 images/min
370 W44°C
✓ Measured
How we measured this. Every result comes from our own pinned, reproducible AI suite, 12 workloads: the Qwen3-4B to Llama-70B LLM ladder (llama.cpp, Q4_K_M), SDXL / Z-Image / FLUX-dev generation, FLUX-Kontext / Qwen-Edit editing, and LTX / Wan video, run first-party on rented hardware with under 0.5% run-to-run variance. Peak VRAM, power draw, temperature and tokens-per-watt are captured per workload. “Won’t fit” rows are real data: where a model exceeds the card’s VRAM at the tested precision we record a hard gate rather than silently dropping to a smaller quant. Measured 2026-07-11 · harness 2.0.0-standalone.

NVIDIA RTX PRO 6000 Blackwell Workstation Edition specifications

ArchitectureBlackwell
CUDA cores24,064
VRAM96GB GDDR7 ECC
Memory bus512-bit
Memory bandwidth1792 GB/s
Boost clock2,617 MHz
TDP600 W
Process4nm (TSMC 4N)
InterfacePCIe 5.0 x16
Release date2025-03-18
Launch MSRP$8,565

Verdict, NVIDIA RTX PRO 6000 Blackwell Workstation Edition on real AI workloads

NVIDIA RTX PRO 6000 Blackwell Workstation Edition scores 53.1/100, #9 of 102. It ran all 12 workloads. Every figure here is our own measurement.

Relative performance: where the NVIDIA RTX PRO 6000 Blackwell Workstation Edition lands

100% = this card, AI & Machine Learning headline metric (AI Score). #1 of 20 desktop cards in this vertical.

GPURelative%AI Score
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
100%53.1
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
85%45.3
NVIDIA RTX PRO 5000 Blackwell
60%31.6
NVIDIA RTX A6000
40%21
NVIDIA RTX PRO 4500 Blackwell
24%12.9

← All AI & Machine Learning GPU rankings

What this card can build

Whole-job timings, composed from our measured per-model results on this card.

WorkflowTimeEnergyBasis
50-image depth pass4 s0.06 Whmeasured
500-image masking run2.6 min12.84 Whmeasured
24-frame storyboard3.6 min16.08 Whall 2 stages measured
200-product catalogue cutout4 min23.56 Whall 2 stages measured
60-second AI short film4.5 min19.26 Whall 3 stages measured
6-panel comic page5.1 min28.55 Whall 3 stages measured
Character sheet, 12 poses6.1 min40.17 Whall 2 stages measured
Full codebase review6.8 min27 Whmeasured
20 long-form articles13.4 min42.96 Whmeasured
40-product photo shoot16.7 min142.82 Whall 2 stages measured
40-product shoot, start to finish17.6 min147.54 Whall 4 stages measured
100-photo restoration batch33.4 min322.75 Whmeasured
100-photo restore and enlarge35.3 min334.41 Whall 2 stages measured