Datacenter platforms · B300 measured first-party · Updated July 2026
HGX B300 is not a server you can buy. It's the GPU baseboard platform NVIDIA sells to server manufacturers, Supermicro, Dell, HPE, who build it into machines with their own CPUs, cooling and networking. That distinction confuses nearly everyone who searches for it, so this page draws the line clearly. Then it does what nobody else does: tells you what one of the eight B300s on that board actually delivers, measured.

288GB of HBM3e at 8 TB/s. We ran it through 21 workloads on 2026-07-12: 47.97 tok/s on Llama 3.3 70B, 9.71 it/s on FLUX.1-dev, 720p Wan video at 2.94 frames/s. The only card on our board of 102 where nothing in the core suite fails to fit.
Best for: Understanding what one B300 actually does before you spec, rent or buy anything built from them.
Almost every question about HGX B300 is really a question about a word. HGX is a platform, not a product. NVIDIA does not sell you an HGX B300 server, because NVIDIA does not make an HGX B300 server. It makes the baseboard, eight B300 Blackwell Ultra GPUs mounted together with the NVLink switching that connects them: and sells that to companies like Supermicro, Dell and HPE, who wrap it in their own CPUs, memory, storage, networking, power delivery and cooling, and sell you the result. DGX B300 is what happens when NVIDIA does that wrapping itself. Same eight GPUs, same fabric, but NVIDIA picks the Xeons, NVIDIA picks the networking, NVIDIA pre-installs the software, and NVIDIA supports it. If you buy an HGX B300 system from Supermicro, you get Supermicro's engineering choices and support contract, usually at a lower price with more configuration freedom. That's the entire distinction. The GPUs don't know which one they're in. Which is exactly what makes the per-GPU number useful, and why we measured it.
One B300 on the HGX board, full LLM ladder, single stream, Q4_K_M
All 11 language models we ran on the B300, including the 6 Blackwell Ultra extras. Every one fit in 288GB with no offload.
On images at BF16: SDXL at 14.6 it/s, FLUX.1-dev at 9.71 it/s, FLUX.1 Kontext at 4.82 it/s, Qwen-Image-Edit at 4.07 it/s. On video: LTX at 31.87 frames/s, Wan 2.2 at 720p at 2.94 frames/s. All 12 core workloads fit, the only card on our board of which that's true. The density question is where HGX gets interesting, and where OEM designs diverge from NVIDIA's.
Supermicro's published HGX B300 rack examples, cooling decides density
| Rack design | B300 GPUs | HBM3e per rack |
|---|---|---|
| Air-cooled example | 32 | 9.2 TB |
| Liquid-cooled example | 64 | 18.4 TB |
Source: Supermicro HGX B300 system datasheet. Cooling architecture, not budget, is what decides how many GPUs you get per rack, the practical consequence of a 1,400W-rated part.
That's the thing about HGX: because you're buying a server rather than a system, every one of these decisions lands on you rather than on NVIDIA. More freedom, more rope.
HGX B300 is the eight-GPU Blackwell Ultra baseboard OEMs build servers around; DGX B300 is NVIDIA's own server built on the same thing. Identical silicon, different assembler. The number no platform spec gives you is per-GPU, so we measured it: 47.97 tok/s on Llama 3.3 70B, 9.71 it/s on FLUX.1-dev, and nothing in our suite too big for 288GB. If you're speccing a deployment, take the 1,400W rating seriously for training and sceptically for inference, we never saw above 1057.7W, and LLM work sat near 350W.
Every number on this page describing a single B300 is our own measurement. We rented a B300 and ran the GPU Battle AI Suite v2 across 21 workloads, the core 12 that every GPU on this site runs, plus 9 Blackwell Ultra extras. LLMs run on llama.cpp (llama-bench) at Q4_K_M with -p 512 -n 128; diffusion and video run on diffusers/ComfyUI at BF16 (SDXL at FP16). Each workload gets a warmup pass plus multiple timed runs (5 for small LLMs, 3 for large models and images, 2 for video). We publish the mean as the result and the minimum as the 1% low. Run-to-run variance across our fleet is under 0.5%. Telemetry is sampled at 1 Hz from nvidia-smi for the duration of every run: power, temperature, utilisation, clocks and peak VRAM. Every wattage, temperature and tokens-per-watt figure on this page is logged draw, not a board rating. The B300 was measured on 2026-07-12 on a CUDA 13 stack (harness 2.1.0-b300-cuda13, driver 580.95.05, torch 2.13.0+cu130), which differs from the CUDA 12.8 stack the rest of our fleet runs. We label it rather than hide it. The caveat that matters most: these are single-GPU, single-stream, batch-size-1 numbers. That is the honest way to measure what one chip does, and it is deliberately not how a datacenter runs a B300. Vendor and MLPerf figures use large batches across many GPUs and will be far higher. Neither is wrong, they answer different questions. Ours answers 'what does one of these actually do'. System-level specifications (GPU counts, memory totals, rack power) come from NVIDIA's published documentation and are cited as such; we have not taken a rack apart. Where NVIDIA's own published system figures disagree with each other, we show the arithmetic rather than pick a side.