48GB · AI Score 31.6/100 · first-party measured on 12 AI workloads
31.6 AI Score ✓ Measured
Every number on this page is first-party: NVIDIA RTX PRO 5000 Blackwell was run on our pinned 12-workload AI suite on 2026-07-12, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) NVIDIA RTX PRO 5000 Blackwell delivers about 208.3 tokens/sec. Stepping up to Qwen3 32B it holds roughly 54.27 tok/s. The full Llama 3.3 70B still runs, at about 26.29 tok/s. For image generation, SDXL runs at 8.73 it/s, and FLUX.1-dev at 1.96 it/s. All 12 workloads fit in 48GB. There is no model in our suite this card has to turn down.
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Qwen3 4B | 298.78 tok/s | 3.1 GB peak131 W47°C2.29 tok/WQ4_K_M | ✓ Measured |
| Llama 3.1 8B | 208.3 tok/s | 4.8 GB peak171 W49°C1.22 tok/WQ4_K_M | ✓ Measured |
| Qwen2.5-Coder 14B | 114.63 tok/s | 8.8 GB peak181 W54°C0.63 tok/WQ4_K_M | ✓ Measured |
| Qwen3 32B | 54.27 tok/s | 18.8 GB peak166 W57°C0.33 tok/WQ4_K_M | ✓ Measured |
| Llama 3.3 70B | 26.29 tok/s | 39.9 GB peak158 W61°C0.17 tok/WQ4_K_M | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Stable Diffusion XL | 17.46 images/min | 14.5 GB peak297 W63°C3.4 s/img | ✓ Measured |
| Z-Image Turbo | 9.08 images/min | 25.7 GB peak300 W67°C6.6 s/img | ✓ Measured |
| FLUX.1 dev | 4.2 images/min | 36.6 GB peak300 W72°C14.3 s/img | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| FLUX.1 Kontext dev | 1.8 images/min | 35.4 GB peak300 W76°C33.2 s/img | ✓ Measured |
| Qwen-Image-Edit | 0.8 images/min | 40.1 GB peak185 W71°C74.9 s/img | ✓ Measured CPU offload |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| LTX-Video (distilled) | 4.58 frames/s | 9.4 GB peak174 W64°C21.2 s/clip | ✓ Measured CPU offload |
| Wan 2.2 5B (720p) | 0.56 frames/s | 37.4 GB peak300 W77°C88.3 s/clip | ✓ Measured |
| Architecture | Blackwell |
| CUDA cores | 14,080 |
| VRAM | 48GB GDDR7 ECC |
| Memory bus | 384-bit |
| Memory bandwidth | 1344 GB/s |
| Boost clock | 2,377 MHz |
| TDP | 300 W |
| Process | 4nm (TSMC 4N) |
| Interface | PCIe 5.0 x16 |
| Release date | 2025-08-01 |
| Launch MSRP | $4,500 |
NVIDIA RTX PRO 5000 Blackwell scores 31.6/100, #16 of 102. It ran all 12 workloads. Every figure here is our own measurement.
100% = this card, AI & Machine Learning headline metric (AI Score). #3 of 20 workstation cards in this vertical.
| GPU | Relative | % | AI Score |
|---|---|---|---|
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 168% | 53.1 | |
| NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition | 138% | 43.5 | |
| NVIDIA RTX PRO 5000 Blackwell | 100% | 31.6 | |
| NVIDIA RTX A6000 | 66% | 21 | |
| NVIDIA RTX PRO 4500 Blackwell | 41% | 12.9 | |
| NVIDIA RTX A5500 | 27% | 8.4 | |
| NVIDIA Quadro RTX 8000 | 26% | 8.2 |
← All AI & Machine Learning GPU rankings
| Transistors | 92,200 million |
| Die size | 750 mm² |
| Process node | 4 nm |
| Fabricated by | TSMC |
| Transistor density | 122.9 million per mm² |
Denser than 96% of the 746 cards we have silicon data for. Density is the clearest measure of what a process node bought: a card that gained it without growing the die got its speed from the fab rather than the architecture.
Silicon figures from Wikipedia (CC BY-SA 4.0). Benchmarks on this page are our own. Compare every chip.
Whole-job timings, composed from our measured per-model results on this card.
| Workflow | Time | Energy | Basis |
|---|---|---|---|
| 24-frame storyboard | 3.4 min | 14.39 Wh | all 2 stages measured |
| 6-panel comic page | 5.6 min | 24.38 Wh | all 3 stages measured |
| 60-second AI short film | 6.7 min | 20.33 Wh | all 3 stages measured |
| Character sheet, 12 poses | 7.6 min | 34.5 Wh | all 2 stages measured |
| Full codebase review | 8.7 min | 26.33 Wh | measured |
| Short social clips | 16.5 min | 78.82 Wh | all 3 stages measured |
| Long-form article batch | 17.8 min | 46.68 Wh | measured |
| Product photo shoot | 25.1 min | 122.39 Wh | all 2 stages measured |
| Photo restoration batch | 55.9 min | 277.59 Wh | measured |
This card is $4,500 to buy. The cheapest listed rate on Vast.ai is $0.593/hour, but that is the floor: we budget $0.712/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 6,324 GPU-hours. Below it you are paying for idle silicon.
| How you would use it | GPU-hours a year | Rental cost a year | Time to break even |
|---|---|---|---|
| 2 hours a day, hobby | 730 | $519 | 8.7 years |
| 8 hours a day, working on it | 2,920 | $2,078 | 2.2 years |
| 24/7, always-on agent | 8,760 | $6,234 | 8.7 months |
At hobby usage this card is very unlikely to pay for itself before it is superseded. Rent it. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.
Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.