Pricing & access · B200 measured first-party · Updated July 2026
An eight-GPU NVIDIA DGX B200 system is reported at roughly $370,000 to $500,000, which works out to somewhere in the region of $46,000-62,000 per B200 once you account for the chassis, CPUs and networking around them. As with every datacenter Blackwell part, that's the wrong number for almost everyone reading this, a B200 is an SXM module rated at 1,000W that doesn't go in any machine you own. The useful answer is that you rent one, which is exactly what we did to measure it across 12 workloads.

192GB HBM3e, also 8 TB/s, 1,000W. Loses to the B300 on 11 of 12 workloads but wins Stable Diffusion XL outright, 23.06 it/s vs 14.6, at 97% utilisation against the B300's 38.5%.
Best for: SDXL-heavy pipelines today, and anyone who wants Blackwell on a stack that's had longer to settle.

288GB HBM3e at 8 TB/s. Wins 11 of our 12 core workloads against the B200, by +4.7% to +80.5%. Then loses Stable Diffusion XL by 37%. We measured both.
Best for: Anyone who needs the memory ceiling gone, and who won't be running FP16 SDXL as their main job.
There's no retail price for a B200 because there's no retail channel. NVIDIA sells Blackwell datacenter silicon into systems and to cloud providers. What gets quoted publicly is the system: a DGX B200, eight B200s, dual Xeon Platinum 8570s, NVSwitch fabric, 14.4 TB/s of GPU-to-GPU bandwidth, up to 4TB of system memory, 30TB of NVMe, 10U: at somewhere around $370,000 to $500,000 depending on vendor and configuration. Divide it out and you're in the region of $46,000-62,000 per GPU, though that's an inference from system pricing rather than a list price. And it's academic. A B200 is a 1,000W SXM module. There is no slot for it in your machine, no connector on your power supply, and nothing you can buy that will move a kilowatt of heat out of a tower case. The question underneath 'how much does a B200 cost' is almost always 'how do I get to use one', and that has a much better answer.
What a few dollars of rented B200 bought us. Measured, 12 workloads
Plus SDXL at 23.06 it/s, FLUX.1-dev at 6.13 it/s, FLUX.1 Kontext at 2.67 it/s, LTX video at 26.8 frames/s and Wan 2.2 720p at 1.88 frames/s. Every one fit in 192GB.
What does the money buy, if not speed? Capacity, mostly, the same answer as every card at this tier. 192GB of HBM3e at 8 TB/s meant all 12 of our core workloads fit with no offload and no quantisation compromise. On our board of 102 AI GPUs, that puts the B200 in a very small club. A 24GB card can't load FLUX.1-dev at BF16 at all. But speed is a smaller part of the story than the price suggests. Single-stream Llama 3.1 8B on a B200 measures 274.41 tok/s. That's a real lead over consumer silicon: but it isn't a 25× lead, because token generation is bound by memory bandwidth, and once a model fits, the gaps compress hard. You are paying for the ability to hold enormous models and to batch enormous concurrency, not for a bigger number on a single chat session. Which is why the honest recommendation is the one we followed ourselves. Rent one, run your actual workload on it, and find out what it does before anybody sends you a quote. We don't publish an hourly rate here because rates drift constantly with supply and region, and a number hardcoded into an evergreen page is wrong by next quarter. The live rate is on the provider's site, where it belongs.
A DGX B200 is roughly $370,000-500,000 for eight GPUs, call it $46,000-62,000 per B200, inferred from system pricing rather than quoted. It's not a price you'll pay, because a 1,000W SXM module isn't a thing an individual buys. Rent it: our full 12-workload measurement run cost a few dollars. What the money buys is 192GB of HBM3e and every workload in our suite fitting without compromise, not a dramatic single-stream speed advantage, 274.41 tok/s on Llama 3.1 8B. And one oddity worth pricing in: on Stable Diffusion XL the B200 is 58% faster than the newer, dearer B300.
Every number describing a single B200 or B300 on this page is our own measurement. Both cards were rented and run through the GPU Battle AI Suite v2, the same 12 core workloads, same models, same settings, same harness. LLMs run on llama.cpp (llama-bench) at Q4_K_M with -p 512 -n 128; diffusion and video run on diffusers/ComfyUI at BF16, with SDXL at FP16. Each workload gets a warmup pass plus multiple timed runs (5 for small LLMs, 3 for large models and images, 2 for video); we publish the mean as the result and the minimum as the 1% low. Telemetry is sampled at 1 Hz from nvidia-smi for the whole run: power, temperature, utilisation, SM clocks and peak VRAM. Every wattage, clock and utilisation figure here is logged, not a board rating. The one asymmetry, and it matters on this page: the B200 was measured on 2026-07-10 on harness 2.0.0 with torch 2.7.0+cu128 (CUDA 12.8). The B300 was measured on 2026-07-12 on harness 2.1.0-b300-cuda13 with torch 2.13.0+cu130 (CUDA 13), because at the time of testing that was the stack the card required. Everything else about the two runs is identical. We label this on every B300 page rather than hide it, and on this page it is the central variable. All figures are single-GPU, single-stream, batch-size-1. That is the honest way to measure what one chip does, and it is not how a datacenter runs these cards. Vendor and MLPerf numbers use large batches across many GPUs and will be far higher. Neither is wrong; they answer different questions.