NVIDIA RTX PRO 4000 Blackwell, AI & Machine Learning Benchmarks & Specs

24GB · AI Score 7.6/100 · first-party measured on 12 AI workloads

7.6 AI Score ✓ Measured

Every number on this page is first-party: NVIDIA RTX PRO 4000 Blackwell was run on our pinned 12-workload AI suite on 2026-07-12, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) NVIDIA RTX PRO 4000 Blackwell delivers about 112.36 tokens/sec. Stepping up to Qwen3 32B it holds roughly 27.8 tok/s. Llama 3.3 70B does not fit. It needs roughly 42GB and this card has 24GB. For image generation, SDXL runs at 4.28 it/s, while FLUX.1-dev won't fit at BF16 (needs ~26GB). 4 of the 12 workloads won't fit on 24GB at the tested precision, Llama 3.3 70B, FLUX.1-dev, FLUX.1 Kontext, Qwen-Image-Edit. We publish those as hard gates rather than quietly dropping to a smaller quant. NVIDIA RTX PRO 4000 Blackwell isn't a retail purchase for most people. It's rented by the hour. You can run this exact card on RunPod.

AI & Machine Learning benchmark results

Text Generation tok/s 18

MiniCPM5 2B271.18
LFM2.5 2.6B254.94
Qwen3 30B A3B185.31
Granite 4.1 3B182.59
Qwen3 4B179.58
Nemotron 3 Nano 4B174.55
Gemma 4 26B A4B134.51
DeepSeek Coder 7B Instruct v1.5117.35
Llama 3.1 8B112.36
Llama 3 8B104.38
Qwen3 8B101.59
Ornith 1.5 9B91.88
WorkloadResultTelemetryData
MiniCPM5 2B271.18 tok/s
81 W46°CQ4_K_M
✓ Measured
LFM2.5 2.6B254.94 tok/s
83 W45°CQ4_K_M
✓ Measured
Granite 4.1 3B182.59 tok/s
90 W52°CQ4_K_M
✓ Measured
Nemotron 3 Nano 4B174.55 tok/s
90 W48°CQ4_K_M
✓ Measured
Qwen3 4B179.58 tok/s
2.7 GB peak96 W43°C1.88 tok/WQ4_K_M
✓ Measured
DeepSeek Coder 7B Instruct v1.5117.35 tok/s
106 W53°CQ4_K_M
✓ Measured
Llama 3 8B104.38 tok/s
103 W53°CQ4_K_M
✓ Measured
Llama 3.1 8B112.36 tok/s
4.9 GB peak112 W47°C1.01 tok/WQ4_K_M
✓ Measured
Qwen3 8B101.59 tok/s
107 W53°CQ4_K_M
✓ Measured
Nemotron Nano 9B v282.32 tok/s
109 W53°CQ4_K_M
✓ Measured
Ornith 1.5 9B91.88 tok/s
105 W51°CQ4_K_M
✓ Measured
Gemma 4 12B64.44 tok/s
108 W53°CQ4_K_M
✓ Measured
Qwen2.5-Coder 14B60.62 tok/s
8.6 GB peak123 W53°C0.49 tok/WQ4_K_M
✓ Measured
Qwen3 14B56.63 tok/s
114 W54°CQ4_K_M
✓ Measured
Gemma 4 26B A4B134.51 tok/s
84 W50°CQ4_K_M
✓ Measured
Qwen3 30B A3B185.31 tok/s
74 W50°CQ4_K_M
✓ Measured
Qwen3 32B27.8 tok/s
18.7 GB peak114 W56°C0.24 tok/WQ4_K_M
✓ Measured
Llama 3.3 70B✕ Won't fit needs ~46 GBVRAM-gated at this precision✓ Measured

Image Generation images/min 7

Stable Diffusion 1.540.91
Sana 1.6B20.12
PixArt-Sigma XL13.44
Stable Diffusion XL8.56
Playground v2.55.54
Z-Image Turbo4.425
WorkloadResultTelemetryData
Stable Diffusion 1.540.91 images/min
140 W51°C
✓ Measured
Sana 1.6B20.12 images/min
145 W60°C
✓ Measured
Stable Diffusion XL8.56 images/min
14.4 GB peak145 W61°C7 s/img
✓ Measured
Playground v2.55.54 images/min
145 W72°C
✓ Measured
PixArt-Sigma XL13.44 images/min
145 W65°C
✓ Measured
Z-Image Turbo4.43 images/min
23.1 GB peak145 W64°C13.6 s/img
✓ Measured
FLUX.1 dev✕ Won't fit needs ~26 GBVRAM-gated at this precision✓ Measured

Image Editing images/min 2

WorkloadResultTelemetryData
FLUX.1 Kontext dev✕ Won't fit needs ~26 GBVRAM-gated at this precision✓ Measured
Qwen-Image-Edit✕ Won't fit needs ~42 GBVRAM-gated at this precision✓ Measured

Video Generation frames/s 2

WorkloadResultTelemetryData
LTX-Video (distilled)3.1 frames/s
9.3 GB peak118 W62°C31.3 s/clip
✓ Measured
CPU offload
Wan 2.2 5B (720p)0.26 frames/s
17.3 GB peak139 W67°C188.5 s/clip
✓ Measured
CPU offload
How we measured this. Every result comes from our own pinned, reproducible AI suite, 12 workloads: the Qwen3-4B to Llama-70B LLM ladder (llama.cpp, Q4_K_M), SDXL / Z-Image / FLUX-dev generation, FLUX-Kontext / Qwen-Edit editing, and LTX / Wan video, run first-party on rented hardware with under 0.5% run-to-run variance. Peak VRAM, power draw, temperature and tokens-per-watt are captured per workload. “Won’t fit” rows are real data: where a model exceeds the card’s VRAM at the tested precision we record a hard gate rather than silently dropping to a smaller quant. Measured 2026-07-12 · harness 2.0.0-standalone.

NVIDIA RTX PRO 4000 Blackwell specifications

ArchitectureBlackwell
CUDA cores8,960
VRAM24GB GDDR7 ECC
Memory bus192-bit
Memory bandwidth672 GB/s
Boost clock2,617 MHz
TDP140 W
Process4nm (TSMC 4N)
InterfacePCIe 5.0 x16
Release date2025-08-01
Launch MSRP$1,500

Verdict, capable, but 24GB sets the ceiling

NVIDIA RTX PRO 4000 Blackwell scores 7.6/100, #34 of 102. It ran 8 of 12; 4 exceeded its 24GB. Every figure here is our own measurement.

Relative performance: where the NVIDIA RTX PRO 4000 Blackwell lands

100% = this card, AI & Machine Learning headline metric (AI Score). #8 of 20 workstation cards in this vertical.

GPURelative%AI Score
NVIDIA RTX A6000
276%21
NVIDIA RTX PRO 4500 Blackwell
170%12.9
NVIDIA RTX A5500
111%8.4
NVIDIA Quadro RTX 8000
108%8.2
NVIDIA RTX PRO 4000 Blackwell
100%7.6
NVIDIA Quadro RTX 6000 (Turing)
97%7.4
NVIDIA RTX A5000
97%7.4
NVIDIA RTX 4500 Ada Generation
78%5.9
NVIDIA RTX A4500
67%5.1

← All AI & Machine Learning GPU rankings

The silicon

Transistors45,600 million
Die size378 mm²
Process node4 nm
Fabricated byTSMC
Transistor density120.6 million per mm²

Denser than 92% of the 746 cards we have silicon data for. Density is the clearest measure of what a process node bought: a card that gained it without growing the die got its speed from the fab rather than the architecture.

Silicon figures from Wikipedia (CC BY-SA 4.0). Benchmarks on this page are our own. Compare every chip.

What this card can build

Whole-job timings, composed from our measured per-model results on this card.

WorkflowTimeEnergyBasis
24-frame storyboard6.4 min14.71 Whall 2 stages measured
60-second AI short film10.2 min20.63 Whall 3 stages measured
Full codebase review16.5 min33.79 Whmeasured
Short social clips34.2 min78.69 Whall 3 stages measured

Can't run: Product photo shoot (needs FLUX.1 Kontext dev), Photo restoration batch (needs FLUX.1 Kontext dev), Restore and enlarge photos (needs FLUX.1 Kontext dev), Character sheet, 12 poses (needs FLUX.1 dev), 6-panel comic page (needs FLUX.1 dev), Long-form article batch (needs Llama 3.3 70B).

Rent or buy?

This card is $1,500 to buy. The cheapest listed rate on Vast.ai is $0.230/hour, but that is the floor: we budget $0.276/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 5,435 GPU-hours. Below it you are paying for idle silicon.

How you would use itGPU-hours a yearRental cost a yearTime to break even
2 hours a day, hobby730$2017.4 years
8 hours a day, working on it2,920$8061.9 years
24/7, always-on agent8,760$2,4187.4 months

At hobby usage this card is very unlikely to pay for itself before it is superseded. Rent it. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.

Rental price

$0.230/hr+12.7% since 2026-10-01low $0.204 · high $0.230

Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.