Datacenter systems · B200 measured first-party · Updated July 2026

What Is the NVIDIA DGX B200?

The NVIDIA DGX B200 is NVIDIA's 10U AI server built on eight B200 Blackwell GPUs, rated at 72 petaFLOPS for training and 144 petaFLOPS for inference, with 1.44TB of total GPU memory and dual Intel Xeon Platinum 8570 CPUs. Every one of those numbers describes the whole box. We rented a single B200 and measured it across our 12-workload suite, so this page can tell you what one eighth of a DGX B200 actually does, 44.54 tok/s on Llama 3.3 70B, since you asked.

Blackwell
NVIDIA B200

NVIDIA B200

192GB HBM3e, also 8 TB/s, 1,000W. Loses to the B300 on 11 of 12 workloads but wins Stable Diffusion XL outright, 23.06 it/s vs 14.6, at 97% utilisation against the B300's 38.5%.

Pros
  • 23.06 it/s on SDXL, beats the B300 by 58%
  • 44.54 tok/s on Llama 3.3 70B (measured)
  • All 12 core workloads fit in 192GB
  • AI Score 78.0/100 · mature CUDA 12.8 stack
Cons
  • Slower than the B300 on the other 11 workloads
  • 192GB vs 288GB
  • 1,000W board rating

Best for: SDXL-heavy pipelines today, and anyone who wants Blackwell on a stack that's had longer to settle.

8×B200
GPUs per DGX B200
10U, dual Xeon Platinum 8570
1.44TB
Total GPU memory
= 180GB per GPU, not 192, see below
72 / 144PF
Training / inference rating
across all eight GPUs
44.54tok/s
What ONE B200 does on Llama 3.3 70B
our measurement, single stream

The DGX B200 is NVIDIA's own finished server built on the HGX B200 platform: eight B200 Blackwell GPUs connected by fifth-generation NVLink through NVSwitch, dual Intel Xeon Platinum 8570 processors with 112 cores, up to 4TB of system memory, 30TB of NVMe, and NVIDIA's full software stack pre-installed. NVIDIA rates it at 72 petaFLOPS of training performance and 144 petaFLOPS of inference, claims 3× the training and 15× the inference performance of the previous generation, and quotes 14.4 TB/s of bidirectional GPU-to-GPU bandwidth. Reported system pricing sits somewhere around $370,000-500,000 depending on configuration and vendor.

As with every DGX, the published numbers describe the chassis. What one B200 does is the number you actually need, and it isn't on the datasheet. So we measured it.

One B200 out of the eight. Measured LLM throughput, single stream, Q4_K_M

Qwen3 4B
317.93 tok/s
Llama 3.1 8B
274.41 tok/s
Qwen2.5-Coder 14B
150.96 tok/s
Qwen3 32B
78.56 tok/s
Llama 3.3 70B
44.54 tok/s

All 12 of our core workloads fit in the B200's 192GB with no offload. Multiply by 8 for a DGX B200 ceiling if you'd run eight independent jobs, roughly 356.3 tok/s aggregate on a 70B. Not if you're running one job across all eight: that's what the NVSwitch fabric is for, and the maths stops being linear.

The other thing worth knowing before you spec one is what the B200 is and isn't fast at, which we can say because we measured it against its successor on identical workloads. Against the B300, the B200 loses almost everywhere: 4.7% slower on Llama 3.1 8B, 7.7% slower on Llama 3.3 70B, 58.4% slower on FLUX.1-dev, 80.5% slower on FLUX.1 Kontext. That's the generational step, and it's real. Then there's Stable Diffusion XL, where the B200 is 58% faster than the B300. That's not a typo. It's the most interesting result in our entire Blackwell dataset, and we wrote it up separately because the telemetry behind it says the B300 was sitting at 38.5% utilisation while the B200 pinned at 97%. If your DGX B200 is destined for an SDXL pipeline, that's a fact worth having.

Our verdict

A DGX B200 is eight B200 Blackwell GPUs in a 10U chassis: 1.44TB of GPU memory, 72 PFLOPS training, 144 PFLOPS inference, 14.4 TB/s of NVLink, and roughly $370,000-500,000. The per-GPU number NVIDIA doesn't publish is the one we measured: 44.54 tok/s on Llama 3.3 70B, 274.41 tok/s on Llama 3.1 8B, 23.06 it/s on SDXL, all single stream, and all 12 of our workloads fitting in 192GB. Check the memory arithmetic before you size models against it, NVIDIA's own 1,440GB figure implies 180GB per GPU, not 192GB.

FAQ

How many GPUs are in a DGX B200?
Eight B200 Blackwell GPUs, connected by fifth-generation NVLink through NVSwitch, alongside dual Intel Xeon Platinum 8570 CPUs with 112 cores total and up to 4TB of system memory, in a 10U chassis. NVIDIA rates the system at 72 petaFLOPS of training and 144 petaFLOPS of inference performance, both totals across all eight GPUs.
How fast is one B200 in a DGX B200?
We measured a single B200 directly: 274.41 tok/s on Llama 3.1 8B, 78.56 tok/s on Qwen3 32B, 44.54 tok/s on Llama 3.3 70B (all Q4_K_M, single stream), 23.06 it/s on SDXL and 6.13 it/s on FLUX.1-dev at BF16. All 12 of our core workloads fit in its 192GB with no offload.
How much GPU memory does a DGX B200 have, 1.44TB or 1.54TB?
NVIDIA's own specification says 1,440GB, which works out to 180GB per GPU across eight. But a standalone B200 is specified at 192GB, and 8 × 192 would be 1,536GB. Some vendor listings describe the DGX/HGX part as a 180GB SKU; others quote 1,536GB. NVIDIA doesn't reconcile the two publicly. The same gap appears on the DGX B300 (2.1TB ÷ 8 = 262.5GB against a 288GB GPU spec), while the GB300 NVL72 reports the full per-GPU amount (20.7TB ÷ 72 = 287.5GB). Whatever the cause, size your models against the system figure rather than the GPU figure.
What is the difference between DGX B200 and DGX B300?
The GPUs. DGX B200 uses B200 (Blackwell); DGX B300 uses B300 (Blackwell Ultra) with more memory per GPU and a higher power budget. On our own measurements of the individual cards, the B300 is 7.7% faster on Llama 3.3 70B, 58.4% faster on FLUX.1-dev and 80.5% faster on FLUX.1 Kontext, but 37% slower on Stable Diffusion XL, which we've written up separately because the cause appears to be software rather than silicon.
What is the difference between DGX B200 and HGX B200?
HGX B200 is the eight-GPU baseboard platform NVIDIA sells to server manufacturers like Supermicro and Dell, who build it into their own machines with their own CPUs, cooling and networking. DGX B200 is NVIDIA's own finished server built on that platform, with NVIDIA's component choices and software stack. Identical GPUs; the difference is who assembled it and who supports it.
Can I rent a B200 instead of buying a DGX B200?
Yes, and it's how we got the numbers on this page. B200 capacity is available by the hour from GPU cloud providers with no commitment, which is a rather different proposition from a system reported at $370,000-500,000. Rates move constantly with supply and provider so we don't publish one here, but our full 12-workload measurement run on a rented B200 cost a few dollars of compute.
Is a DGX B200 worth it over renting?
That depends entirely on utilisation, and it's worth doing the arithmetic honestly. Owned hardware only pays back when it's kept busy. Our single-GPU telemetry across the Blackwell generation consistently shows single-stream inference leaving most of these chips idle: a B200 running SDXL sits at 97% utilisation, but LLM token generation is bandwidth-bound and uses a fraction of the compute. If your workload is batched, concurrent production serving, a DGX earns its price. If it's exploratory or bursty, renting the same silicon by the hour is hard to argue against.

How we test

Every number describing a single B200 or B300 on this page is our own measurement. Both cards were rented and run through the GPU Battle AI Suite v2, the same 12 core workloads, same models, same settings, same harness. LLMs run on llama.cpp (llama-bench) at Q4_K_M with -p 512 -n 128; diffusion and video run on diffusers/ComfyUI at BF16, with SDXL at FP16. Each workload gets a warmup pass plus multiple timed runs (5 for small LLMs, 3 for large models and images, 2 for video); we publish the mean as the result and the minimum as the 1% low. Telemetry is sampled at 1 Hz from nvidia-smi for the whole run: power, temperature, utilisation, SM clocks and peak VRAM. Every wattage, clock and utilisation figure here is logged, not a board rating. The one asymmetry, and it matters on this page: the B200 was measured on 2026-07-10 on harness 2.0.0 with torch 2.7.0+cu128 (CUDA 12.8). The B300 was measured on 2026-07-12 on harness 2.1.0-b300-cuda13 with torch 2.13.0+cu130 (CUDA 13), because at the time of testing that was the stack the card required. Everything else about the two runs is identical. We label this on every B300 page rather than hide it, and on this page it is the central variable. All figures are single-GPU, single-stream, batch-size-1. That is the honest way to measure what one chip does, and it is not how a datacenter runs these cards. Vendor and MLPerf numbers use large batches across many GPUs and will be far higher. Neither is wrong; they answer different questions.