24GB · AI Score 7.6/100 · first-party measured on 12 AI workloads
7.6 AI Score ✓ Measured
Every number on this page is first-party: NVIDIA RTX PRO 4000 Blackwell was run on our pinned 12-workload AI suite on 2026-07-12, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) NVIDIA RTX PRO 4000 Blackwell delivers about 112.36 tokens/sec. Stepping up to Qwen3 32B it holds roughly 27.8 tok/s. Llama 3.3 70B does not fit. It needs roughly 42GB and this card has 24GB. For image generation, SDXL runs at 4.28 it/s, while FLUX.1-dev won't fit at BF16 (needs ~26GB). 4 of the 12 workloads won't fit on 24GB at the tested precision, Llama 3.3 70B, FLUX.1-dev, FLUX.1 Kontext, Qwen-Image-Edit. We publish those as hard gates rather than quietly dropping to a smaller quant. NVIDIA RTX PRO 4000 Blackwell isn't a retail purchase for most people. It's rented by the hour. You can run this exact card on RunPod.
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| MiniCPM5 2B | 271.18 tok/s | 81 W46°CQ4_K_M | ✓ Measured |
| LFM2.5 2.6B | 254.94 tok/s | 83 W45°CQ4_K_M | ✓ Measured |
| Granite 4.1 3B | 182.59 tok/s | 90 W52°CQ4_K_M | ✓ Measured |
| Nemotron 3 Nano 4B | 174.55 tok/s | 90 W48°CQ4_K_M | ✓ Measured |
| Qwen3 4B | 179.58 tok/s | 2.7 GB peak96 W43°C1.88 tok/WQ4_K_M | ✓ Measured |
| DeepSeek Coder 7B Instruct v1.5 | 117.35 tok/s | 106 W53°CQ4_K_M | ✓ Measured |
| Llama 3 8B | 104.38 tok/s | 103 W53°CQ4_K_M | ✓ Measured |
| Llama 3.1 8B | 112.36 tok/s | 4.9 GB peak112 W47°C1.01 tok/WQ4_K_M | ✓ Measured |
| Qwen3 8B | 101.59 tok/s | 107 W53°CQ4_K_M | ✓ Measured |
| Nemotron Nano 9B v2 | 82.32 tok/s | 109 W53°CQ4_K_M | ✓ Measured |
| Ornith 1.5 9B | 91.88 tok/s | 105 W51°CQ4_K_M | ✓ Measured |
| Gemma 4 12B | 64.44 tok/s | 108 W53°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder 14B | 60.62 tok/s | 8.6 GB peak123 W53°C0.49 tok/WQ4_K_M | ✓ Measured |
| Qwen3 14B | 56.63 tok/s | 114 W54°CQ4_K_M | ✓ Measured |
| Gemma 4 26B A4B | 134.51 tok/s | 84 W50°CQ4_K_M | ✓ Measured |
| Qwen3 30B A3B | 185.31 tok/s | 74 W50°CQ4_K_M | ✓ Measured |
| Qwen3 32B | 27.8 tok/s | 18.7 GB peak114 W56°C0.24 tok/WQ4_K_M | ✓ Measured |
| Llama 3.3 70B | ✕ Won't fit needs ~46 GB | VRAM-gated at this precision | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Stable Diffusion 1.5 | 40.91 images/min | 140 W51°C | ✓ Measured |
| Sana 1.6B | 20.12 images/min | 145 W60°C | ✓ Measured |
| Stable Diffusion XL | 8.56 images/min | 14.4 GB peak145 W61°C7 s/img | ✓ Measured |
| Playground v2.5 | 5.54 images/min | 145 W72°C | ✓ Measured |
| PixArt-Sigma XL | 13.44 images/min | 145 W65°C | ✓ Measured |
| Z-Image Turbo | 4.43 images/min | 23.1 GB peak145 W64°C13.6 s/img | ✓ Measured |
| FLUX.1 dev | ✕ Won't fit needs ~26 GB | VRAM-gated at this precision | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| FLUX.1 Kontext dev | ✕ Won't fit needs ~26 GB | VRAM-gated at this precision | ✓ Measured |
| Qwen-Image-Edit | ✕ Won't fit needs ~42 GB | VRAM-gated at this precision | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| LTX-Video (distilled) | 3.1 frames/s | 9.3 GB peak118 W62°C31.3 s/clip | ✓ Measured CPU offload |
| Wan 2.2 5B (720p) | 0.26 frames/s | 17.3 GB peak139 W67°C188.5 s/clip | ✓ Measured CPU offload |
| Architecture | Blackwell |
| CUDA cores | 8,960 |
| VRAM | 24GB GDDR7 ECC |
| Memory bus | 192-bit |
| Memory bandwidth | 672 GB/s |
| Boost clock | 2,617 MHz |
| TDP | 140 W |
| Process | 4nm (TSMC 4N) |
| Interface | PCIe 5.0 x16 |
| Release date | 2025-08-01 |
| Launch MSRP | $1,500 |
NVIDIA RTX PRO 4000 Blackwell scores 7.6/100, #34 of 102. It ran 8 of 12; 4 exceeded its 24GB. Every figure here is our own measurement.
100% = this card, AI & Machine Learning headline metric (AI Score). #8 of 20 workstation cards in this vertical.
| GPU | Relative | % | AI Score |
|---|---|---|---|
| NVIDIA RTX A6000 | 276% | 21 | |
| NVIDIA RTX PRO 4500 Blackwell | 170% | 12.9 | |
| NVIDIA RTX A5500 | 111% | 8.4 | |
| NVIDIA Quadro RTX 8000 | 108% | 8.2 | |
| NVIDIA RTX PRO 4000 Blackwell | 100% | 7.6 | |
| NVIDIA Quadro RTX 6000 (Turing) | 97% | 7.4 | |
| NVIDIA RTX A5000 | 97% | 7.4 | |
| NVIDIA RTX 4500 Ada Generation | 78% | 5.9 | |
| NVIDIA RTX A4500 | 67% | 5.1 |
← All AI & Machine Learning GPU rankings
| Transistors | 45,600 million |
| Die size | 378 mm² |
| Process node | 4 nm |
| Fabricated by | TSMC |
| Transistor density | 120.6 million per mm² |
Denser than 92% of the 746 cards we have silicon data for. Density is the clearest measure of what a process node bought: a card that gained it without growing the die got its speed from the fab rather than the architecture.
Silicon figures from Wikipedia (CC BY-SA 4.0). Benchmarks on this page are our own. Compare every chip.
Whole-job timings, composed from our measured per-model results on this card.
| Workflow | Time | Energy | Basis |
|---|---|---|---|
| 24-frame storyboard | 6.4 min | 14.71 Wh | all 2 stages measured |
| 60-second AI short film | 10.2 min | 20.63 Wh | all 3 stages measured |
| Full codebase review | 16.5 min | 33.79 Wh | measured |
| Short social clips | 34.2 min | 78.69 Wh | all 3 stages measured |
Can't run: Product photo shoot (needs FLUX.1 Kontext dev), Photo restoration batch (needs FLUX.1 Kontext dev), Restore and enlarge photos (needs FLUX.1 Kontext dev), Character sheet, 12 poses (needs FLUX.1 dev), 6-panel comic page (needs FLUX.1 dev), Long-form article batch (needs Llama 3.3 70B).
This card is $1,500 to buy. The cheapest listed rate on Vast.ai is $0.230/hour, but that is the floor: we budget $0.276/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 5,435 GPU-hours. Below it you are paying for idle silicon.
| How you would use it | GPU-hours a year | Rental cost a year | Time to break even |
|---|---|---|---|
| 2 hours a day, hobby | 730 | $201 | 7.4 years |
| 8 hours a day, working on it | 2,920 | $806 | 1.9 years |
| 24/7, always-on agent | 8,760 | $2,418 | 7.4 months |
At hobby usage this card is very unlikely to pay for itself before it is superseded. Rent it. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.
Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.