Datacenter platforms · B300 measured first-party · Updated July 2026

What Is NVIDIA HGX B300?

HGX B300 is not a server you can buy. It's the GPU baseboard platform NVIDIA sells to server manufacturers, Supermicro, Dell, HPE, who build it into machines with their own CPUs, cooling and networking. That distinction confuses nearly everyone who searches for it, so this page draws the line clearly. Then it does what nobody else does: tells you what one of the eight B300s on that board actually delivers, measured.

The GPU we measured
NVIDIA B300

NVIDIA B300

288GB of HBM3e at 8 TB/s. We ran it through 21 workloads on 2026-07-12: 47.97 tok/s on Llama 3.3 70B, 9.71 it/s on FLUX.1-dev, 720p Wan video at 2.94 frames/s. The only card on our board of 102 where nothing in the core suite fails to fit.

Pros
  • 47.97 tok/s on Llama 3.3 70B (measured, single stream)
  • 288GB HBM3e, 12/12 core workloads fit, zero offload
  • 8 TB/s bandwidth, the highest on our board
  • AI Score 93.8/100, the highest score on our board
Cons
  • Roughly $53,000 per GPU; no retail channel exists
  • 1,400W rating, but we never measured above 1057.7W
  • Single-stream LLM work runs it at 12.2-40.3% utilisation
  • Measured on a CUDA 13 stack, unlike the rest of our fleet

Best for: Understanding what one B300 actually does before you spec, rent or buy anything built from them.

8×B300
GPUs per HGX B300 baseboard
same count as DGX B300
14.4TB/s
NVLink Switch bandwidth
NVIDIA platform spec
2.3TB
HBM3e per board
8 × 288GB
287.23tok/s
What ONE B300 does on Llama 3.1 8B
our measurement

Almost every question about HGX B300 is really a question about a word. HGX is a platform, not a product. NVIDIA does not sell you an HGX B300 server, because NVIDIA does not make an HGX B300 server. It makes the baseboard, eight B300 Blackwell Ultra GPUs mounted together with the NVLink switching that connects them: and sells that to companies like Supermicro, Dell and HPE, who wrap it in their own CPUs, memory, storage, networking, power delivery and cooling, and sell you the result. DGX B300 is what happens when NVIDIA does that wrapping itself. Same eight GPUs, same fabric, but NVIDIA picks the Xeons, NVIDIA picks the networking, NVIDIA pre-installs the software, and NVIDIA supports it. If you buy an HGX B300 system from Supermicro, you get Supermicro's engineering choices and support contract, usually at a lower price with more configuration freedom. That's the entire distinction. The GPUs don't know which one they're in. Which is exactly what makes the per-GPU number useful, and why we measured it.

One B300 on the HGX board, full LLM ladder, single stream, Q4_K_M

Qwen3 4B
333.34 tok/s
Llama 3.1 8B
287.23 tok/s
Qwen2.5-Coder 14B
158.49 tok/s
Qwen3 32B
83.68 tok/s
Llama 3.3 70B
47.97 tok/s
GPT-OSS 20B
346.76 tok/s
Qwen3 14B
168.03 tok/s
Phi-4 14B
178.74 tok/s
Gemma 3 27B
93.31 tok/s
Mistral Small 24B
120.55 tok/s
DeepSeek-R1 Distill 8B
289.05 tok/s

All 11 language models we ran on the B300, including the 6 Blackwell Ultra extras. Every one fit in 288GB with no offload.

On images at BF16: SDXL at 14.6 it/s, FLUX.1-dev at 9.71 it/s, FLUX.1 Kontext at 4.82 it/s, Qwen-Image-Edit at 4.07 it/s. On video: LTX at 31.87 frames/s, Wan 2.2 at 720p at 2.94 frames/s. All 12 core workloads fit, the only card on our board of which that's true. The density question is where HGX gets interesting, and where OEM designs diverge from NVIDIA's.

Supermicro's published HGX B300 rack examples, cooling decides density

Rack designB300 GPUsHBM3e per rack
Air-cooled example329.2 TB
Liquid-cooled example6418.4 TB

Source: Supermicro HGX B300 system datasheet. Cooling architecture, not budget, is what decides how many GPUs you get per rack, the practical consequence of a 1,400W-rated part.

That's the thing about HGX: because you're buying a server rather than a system, every one of these decisions lands on you rather than on NVIDIA. More freedom, more rope.

Our verdict

HGX B300 is the eight-GPU Blackwell Ultra baseboard OEMs build servers around; DGX B300 is NVIDIA's own server built on the same thing. Identical silicon, different assembler. The number no platform spec gives you is per-GPU, so we measured it: 47.97 tok/s on Llama 3.3 70B, 9.71 it/s on FLUX.1-dev, and nothing in our suite too big for 288GB. If you're speccing a deployment, take the 1,400W rating seriously for training and sceptically for inference, we never saw above 1057.7W, and LLM work sat near 350W.

FAQ

What is the difference between HGX B300 and DGX B300?
HGX B300 is the platform; DGX B300 is a product built on it. NVIDIA sells the HGX B300 baseboard, eight B300 GPUs plus the NVLink switching between them, to server manufacturers, who add their own CPUs, memory, storage, networking and cooling. DGX B300 is NVIDIA's own finished machine with NVIDIA's component choices and software stack. The GPUs are identical. What differs is who assembled everything around them, and who you call when it breaks.
How many GPUs are on an HGX B300 board?
Eight, in the standard configuration, the same as a DGX B300, since DGX B300 is built on the HGX platform. That's roughly 2.3TB of raw HBM3e per board. NVIDIA quotes 'over 2TB' of high-speed memory and up to 14.4 TB/s of NVLink Switch bandwidth for the platform.
How fast is a single B300 GPU on an HGX B300?
We measured one directly. Single stream, Q4_K_M: 287.23 tok/s on Llama 3.1 8B, 158.49 tok/s on Qwen2.5-Coder 14B, 83.68 tok/s on Qwen3 32B, 47.97 tok/s on Llama 3.3 70B. At BF16: 14.6 it/s on SDXL, 9.71 it/s on FLUX.1-dev, 4.82 it/s on FLUX.1 Kontext. The silicon is identical whether it arrives on an HGX board in a Supermicro chassis or inside a DGX, so these apply to both.
Who makes HGX B300 servers?
The major server OEMs. Supermicro publishes HGX B300 designs in both air-cooled and liquid-cooled variants, and Dell, HPE and others build on the platform too. Supermicro's rack examples show the density involved: an air-cooled rack lists 32 B300 GPUs and 9.2TB of HBM3e, a liquid-cooled one lists 64 GPUs and 18.4TB. Which you can deploy depends less on budget than on what your facility can cool.
Do I need liquid cooling for HGX B300?
It depends on density, and the rating overstates the problem for many workloads. Each B300 carries a 1,400W rating, and OEM designs roughly double GPUs per rack moving from air to liquid. But our telemetry across 21 workloads never saw a B300 exceed 1057.7W, and language-model work averaged 286.2-391.6W. Sizing purely off 1,400W will overbuild for inference-heavy deployments; sizing off our numbers will underbuild for sustained FP4 training. Neither substitutes for measuring your own workload.
Is HGX B300 better than HGX B200?
For memory-bound work, meaningfully. B200 to B300 is 192GB to 288GB per GPU plus a bandwidth increase to 8 TB/s, and NVIDIA's claimed 1.5× dense FP4 and 2× attention performance. NVIDIA states HGX B300 delivers up to 2.6× the training performance of the prior generation on large language models such as DeepSeek R1. If memory capacity is what's limiting you, that extra 96GB per GPU is the whole argument. If it isn't, B200 is the cheaper answer.

How we test

Every number on this page describing a single B300 is our own measurement. We rented a B300 and ran the GPU Battle AI Suite v2 across 21 workloads, the core 12 that every GPU on this site runs, plus 9 Blackwell Ultra extras. LLMs run on llama.cpp (llama-bench) at Q4_K_M with -p 512 -n 128; diffusion and video run on diffusers/ComfyUI at BF16 (SDXL at FP16). Each workload gets a warmup pass plus multiple timed runs (5 for small LLMs, 3 for large models and images, 2 for video). We publish the mean as the result and the minimum as the 1% low. Run-to-run variance across our fleet is under 0.5%. Telemetry is sampled at 1 Hz from nvidia-smi for the duration of every run: power, temperature, utilisation, clocks and peak VRAM. Every wattage, temperature and tokens-per-watt figure on this page is logged draw, not a board rating. The B300 was measured on 2026-07-12 on a CUDA 13 stack (harness 2.1.0-b300-cuda13, driver 580.95.05, torch 2.13.0+cu130), which differs from the CUDA 12.8 stack the rest of our fleet runs. We label it rather than hide it. The caveat that matters most: these are single-GPU, single-stream, batch-size-1 numbers. That is the honest way to measure what one chip does, and it is deliberately not how a datacenter runs a B300. Vendor and MLPerf figures use large batches across many GPUs and will be far higher. Neither is wrong, they answer different questions. Ours answers 'what does one of these actually do'. System-level specifications (GPU counts, memory totals, rack power) come from NVIDIA's published documentation and are cited as such; we have not taken a rack apart. Where NVIDIA's own published system figures disagree with each other, we show the arithmetic rather than pick a side.