Pricing & access · B200 measured first-party · Updated July 2026

How Much Does an NVIDIA B200 Cost?

An eight-GPU NVIDIA DGX B200 system is reported at roughly $370,000 to $500,000, which works out to somewhere in the region of $46,000-62,000 per B200 once you account for the chassis, CPUs and networking around them. As with every datacenter Blackwell part, that's the wrong number for almost everyone reading this, a B200 is an SXM module rated at 1,000W that doesn't go in any machine you own. The useful answer is that you rent one, which is exactly what we did to measure it across 12 workloads.

Blackwell
NVIDIA B200

NVIDIA B200

192GB HBM3e, also 8 TB/s, 1,000W. Loses to the B300 on 11 of 12 workloads but wins Stable Diffusion XL outright, 23.06 it/s vs 14.6, at 97% utilisation against the B300's 38.5%.

Pros
  • 23.06 it/s on SDXL, beats the B300 by 58%
  • 44.54 tok/s on Llama 3.3 70B (measured)
  • All 12 core workloads fit in 192GB
  • AI Score 78.0/100 · mature CUDA 12.8 stack
Cons
  • Slower than the B300 on the other 11 workloads
  • 192GB vs 288GB
  • 1,000W board rating

Best for: SDXL-heavy pipelines today, and anyone who wants Blackwell on a stack that's had longer to settle.

Blackwell Ultra
NVIDIA B300

NVIDIA B300

288GB HBM3e at 8 TB/s. Wins 11 of our 12 core workloads against the B200, by +4.7% to +80.5%. Then loses Stable Diffusion XL by 37%. We measured both.

Pros
  • 47.97 tok/s on Llama 3.3 70B (measured)
  • 9.71 it/s on FLUX.1-dev, 58% faster than a B200
  • 288GB, nothing in our suite fails to fit
  • AI Score 93.8/100
Cons
  • Loses SDXL to the B200: 14.6 vs 23.06 it/s
  • Only 38.5% utilised on that run, the chip sat idle
  • Roughly $53,000 per GPU
  • Measured on a CUDA 13 stack, unlike our CUDA 12.8 fleet

Best for: Anyone who needs the memory ceiling gone, and who won't be running FP16 SDXL as their main job.

$370-500k
DGX B200 system (8 GPUs)
reported system pricing
1,000W
Per-GPU board rating
SXM module, no consumer channel
44.54tok/s
Llama 3.3 70B, one B200
measured, single stream
12 / 12
Workloads that fit in 192GB
no offload, no compromise

There's no retail price for a B200 because there's no retail channel. NVIDIA sells Blackwell datacenter silicon into systems and to cloud providers. What gets quoted publicly is the system: a DGX B200, eight B200s, dual Xeon Platinum 8570s, NVSwitch fabric, 14.4 TB/s of GPU-to-GPU bandwidth, up to 4TB of system memory, 30TB of NVMe, 10U: at somewhere around $370,000 to $500,000 depending on vendor and configuration. Divide it out and you're in the region of $46,000-62,000 per GPU, though that's an inference from system pricing rather than a list price. And it's academic. A B200 is a 1,000W SXM module. There is no slot for it in your machine, no connector on your power supply, and nothing you can buy that will move a kilowatt of heat out of a tower case. The question underneath 'how much does a B200 cost' is almost always 'how do I get to use one', and that has a much better answer.

What a few dollars of rented B200 bought us. Measured, 12 workloads

Qwen3 4B
317.93 tok/s
Llama 3.1 8B
274.41 tok/s
Qwen2.5-Coder 14B
150.96 tok/s
Qwen3 32B
78.56 tok/s
Llama 3.3 70B
44.54 tok/s

Plus SDXL at 23.06 it/s, FLUX.1-dev at 6.13 it/s, FLUX.1 Kontext at 2.67 it/s, LTX video at 26.8 frames/s and Wan 2.2 720p at 1.88 frames/s. Every one fit in 192GB.

What does the money buy, if not speed? Capacity, mostly, the same answer as every card at this tier. 192GB of HBM3e at 8 TB/s meant all 12 of our core workloads fit with no offload and no quantisation compromise. On our board of 102 AI GPUs, that puts the B200 in a very small club. A 24GB card can't load FLUX.1-dev at BF16 at all. But speed is a smaller part of the story than the price suggests. Single-stream Llama 3.1 8B on a B200 measures 274.41 tok/s. That's a real lead over consumer silicon: but it isn't a 25× lead, because token generation is bound by memory bandwidth, and once a model fits, the gaps compress hard. You are paying for the ability to hold enormous models and to batch enormous concurrency, not for a bigger number on a single chat session. Which is why the honest recommendation is the one we followed ourselves. Rent one, run your actual workload on it, and find out what it does before anybody sends you a quote. We don't publish an hourly rate here because rates drift constantly with supply and region, and a number hardcoded into an evergreen page is wrong by next quarter. The live rate is on the provider's site, where it belongs.

Our verdict

A DGX B200 is roughly $370,000-500,000 for eight GPUs, call it $46,000-62,000 per B200, inferred from system pricing rather than quoted. It's not a price you'll pay, because a 1,000W SXM module isn't a thing an individual buys. Rent it: our full 12-workload measurement run cost a few dollars. What the money buys is 192GB of HBM3e and every workload in our suite fitting without compromise, not a dramatic single-stream speed advantage, 274.41 tok/s on Llama 3.1 8B. And one oddity worth pricing in: on Stable Diffusion XL the B200 is 58% faster than the newer, dearer B300.

FAQ

How much does one NVIDIA B200 GPU cost?
There's no list price, B200s are sold into systems and to cloud providers, not through retail. Working back from reported DGX B200 system pricing of roughly $370,000 to $500,000 for eight GPUs plus chassis, CPUs, networking and storage, the implied per-GPU figure lands somewhere around $46,000-62,000. Treat that as an inference, not a quote.
How much does a DGX B200 cost?
Reported pricing runs roughly $370,000 to $500,000 depending on vendor and configuration. That buys eight B200 GPUs, dual Intel Xeon Platinum 8570 processors with 112 cores, NVSwitch fabric with 14.4 TB/s of bidirectional GPU-to-GPU bandwidth, up to 4TB of system memory, 30TB of NVMe storage, and NVIDIA's software stack and support, in a 10U chassis rated at 72 petaFLOPS training and 144 petaFLOPS inference.
Is it cheaper to rent a B200 or buy one?
For anything short of continuous, well-utilised production, renting wins comfortably. We measured a B200 across all 12 workloads in our suite, every LLM up to Llama 3.3 70B, five image models, two video models, with full telemetry, for a few dollars of rented compute. The break-even for ownership requires keeping the card genuinely busy, and single-stream inference doesn't come close: token generation is bandwidth-bound and leaves most of the chip idle.
Is a B200 or a B300 better value?
Depends entirely on the workload, and we have both measured. The B300 wins 11 of our 12 workloads, +7.7% on Llama 3.3 70B, +58.4% on FLUX.1-dev, +80.5% on FLUX.1 Kontext. And has 288GB against the B200's 192GB. But on Stable Diffusion XL the B200 wins by 58% (23.06 vs 14.6 it/s), and it's a cheaper, lower-power part. If your work is SDXL, the B200 is the better buy today. If it's modern diffusion, long context, or anything needing more than 192GB, the B300 earns the premium.
What can a B200 run that a consumer GPU can't?
All 12 of our core workloads fit in its 192GB with no offload, including Llama 3.3 70B at 44.54 tok/s, which needs roughly 42GB and won't load on any consumer card, and FLUX.1-dev at BF16, which needs ~26GB and turns down a 24GB card outright. That's the product: not speed, but the absence of the ceiling. Our whole board of 102 GPUs is mostly a catalogue of what each card has to refuse.
Why is there no B200 hourly rate on this page?
Because it would be wrong within a quarter. Rental rates move constantly with supply, provider, region and commitment level, and hardcoding one into an evergreen page just creates a maintenance treadmill and misleads whoever reads it six months later. The live rate lives on the provider's site. What we can tell you concretely is that our entire 12-workload measurement campaign on a rented B200 cost a few dollars.

How we test

Every number describing a single B200 or B300 on this page is our own measurement. Both cards were rented and run through the GPU Battle AI Suite v2, the same 12 core workloads, same models, same settings, same harness. LLMs run on llama.cpp (llama-bench) at Q4_K_M with -p 512 -n 128; diffusion and video run on diffusers/ComfyUI at BF16, with SDXL at FP16. Each workload gets a warmup pass plus multiple timed runs (5 for small LLMs, 3 for large models and images, 2 for video); we publish the mean as the result and the minimum as the 1% low. Telemetry is sampled at 1 Hz from nvidia-smi for the whole run: power, temperature, utilisation, SM clocks and peak VRAM. Every wattage, clock and utilisation figure here is logged, not a board rating. The one asymmetry, and it matters on this page: the B200 was measured on 2026-07-10 on harness 2.0.0 with torch 2.7.0+cu128 (CUDA 12.8). The B300 was measured on 2026-07-12 on harness 2.1.0-b300-cuda13 with torch 2.13.0+cu130 (CUDA 13), because at the time of testing that was the stack the card required. Everything else about the two runs is identical. We label this on every B300 page rather than hide it, and on this page it is the central variable. All figures are single-GPU, single-stream, batch-size-1. That is the honest way to measure what one chip does, and it is not how a datacenter runs these cards. Vendor and MLPerf numbers use large batches across many GPUs and will be far higher. Neither is wrong; they answer different questions.