Measured telemetry · 21 workloads logged at 1 Hz · Updated July 2026

How Much Power Does an NVIDIA B300 Actually Use?

The NVIDIA B300 carries a 1,400W board rating. Across 21 workloads we measured with 1 Hz telemetry, it never once drew more than 1057.7W, and on language-model inference it averaged between 286.2W and 391.6W, under 30% of its rating, at utilisation as low as 12.2%. That gap isn't a rounding error. It's the difference between what a Blackwell Ultra is built to do and what most people will actually ask it to do, and as far as we can tell nobody else has published it.

The GPU we measured
NVIDIA B300

NVIDIA B300

288GB of HBM3e at 8 TB/s. We ran it through 21 workloads on 2026-07-12: 47.97 tok/s on Llama 3.3 70B, 9.71 it/s on FLUX.1-dev, 720p Wan video at 2.94 frames/s. The only card on our board of 102 where nothing in the core suite fails to fit.

Pros
  • 47.97 tok/s on Llama 3.3 70B (measured, single stream)
  • 288GB HBM3e, 12/12 core workloads fit, zero offload
  • 8 TB/s bandwidth, the highest on our board
  • AI Score 93.8/100, the highest score on our board
Cons
  • Roughly $53,000 per GPU; no retail channel exists
  • 1,400W rating, but we never measured above 1057.7W
  • Single-stream LLM work runs it at 12.2-40.3% utilisation
  • Measured on a CUDA 13 stack, unlike the rest of our fleet

Best for: Understanding what one B300 actually does before you spec, rent or buy anything built from them.

1,400W
Board rating
the spec sheet number
1,057.7W
Highest we ever measured
76% of rating, on FLUX.1 Kontext
377.8W
Average on Llama 3.3 70B
27% of rating
57°C
Peak temp, all 21 workloads
thermals were a non-event

Every specification sheet for the NVIDIA B300 says 1,400W. We logged power at 1 Hz through 21 workloads and never saw it above 1057.7W. On the workloads most people will actually run, it averaged under 400W.

That's a big enough gap to be worth writing down carefully, because it changes what a B300 is for.

Every language model we ran, average power draw against a 1,400W rating

Qwen3 4B
294.5 W
Llama 3.1 8B
338.3 W
Qwen2.5-Coder 14B
367.2 W
Qwen3 32B
345.9 W
Llama 3.3 70B
377.8 W
GPT-OSS 20B
286.2 W
Qwen3 14B
365.5 W
Phi-4 14B
391.6 W
Gemma 3 27B
352.5 W
Mistral Small 24B
352.1 W
DeepSeek-R1 Distill 8B
348.4 W

1,400W board rating

Every language model, from 4 billion parameters to 70 billion, landed between 286.2W and 391.6W. The full Llama 3.3 70B, the heaviest LLM in our suite, drew 27% of the card's rating.

Now the other end. Diffusion is a completely different workload, and the B300 responds like a different card.

Full telemetry, all 21 workloads, sorted by power draw

FLUX.1 Kontext dev1029.5
Qwen-Image-Edit1019.6
Wan 2.2 5B (720p)1012.6
FLUX.1 dev988.8
Qwen-Image944.6
Z-Image Turbo931.7
FLUX.1 schnell890.2
LTX-Video (distilled)715.6
Stable Diffusion XL488.4
Phi-4 14B391.6
Llama 3.3 70B377.8
Qwen2.5-Coder 14B367.2
Qwen3 14B365.5
Stable Diffusion 1.5360.5
Gemma 3 27B352.5
Mistral Small 24B352.1
DeepSeek-R1 Distill 8B348.4
Qwen3 32B345.9
Llama 3.1 8B338.3
Qwen3 4B294.5
GPT-OSS 20B286.2
WorkloadAvg WPeak W% of 1,400WUtil %Peak °C
FLUX.1 Kontext dev1029.51057.776%9956
Qwen-Image-Edit1019.6105175%96.955
Wan 2.2 5B (720p)1012.61054.675%95.957
FLUX.1 dev988.81043.675%9854
Qwen-Image944.6977.570%92.250
Z-Image Turbo931.7937.367%9550
FLUX.1 schnell890.2890.264%9848
LTX-Video (distilled)715.61004.872%65.851
Stable Diffusion XL488.4545.239%38.541
Phi-4 14B391.6692.549%37.644
Llama 3.3 70B377.8780.956%31.346
Qwen2.5-Coder 14B367.2622.544%37.442
Qwen3 14B365.5643.546%40.343
Stable Diffusion 1.5360.5367.526%3238
Gemma 3 27B352.5661.547%31.843
Mistral Small 24B352.1705.650%32.844
DeepSeek-R1 Distill 8B348.4617.144%40.142
Qwen3 32B345.9690.249%25.944
Llama 3.1 8B338.3615.544%3741
Qwen3 4B294.5491.935%33.839
GPT-OSS 20B286.2520.837%12.240

Logged at 1 Hz from nvidia-smi across the whole run. Same silicon, same node, same session, a 3.6× swing in power draw between GPT-OSS 20B and FLUX.1 Kontext. The workload decides, not the card.

Tokens per watt. Measured, by model

GPT-OSS 20B
1.216 tok/W
Qwen3 4B
1.133 tok/W
Llama 3.1 8B
0.851 tok/W
DeepSeek-R1 Distill 8B
0.831 tok/W
Qwen3 14B
0.46 tok/W
Phi-4 14B
0.457 tok/W
Qwen2.5-Coder 14B
0.432 tok/W
Mistral Small 24B
0.343 tok/W
Gemma 3 27B
0.265 tok/W
Qwen3 32B
0.242 tok/W
Llama 3.3 70B
0.127 tok/W

From 1.216 tok/W on GPT-OSS 20B down to 0.127 tok/W on Llama 3.3 70B.

Thermals were a non-event. Peak temperature across all 21 workloads was 57°C. Language models sat between 39°C and 46°C. Even the kilowatt-class diffusion runs topped out in the mid-50s. That says more about the cooling on the node we rented than about the chip, but it does mean nothing we ran was thermally limited. The constraints on a B300 are memory capacity and workload shape, not heat.

What should you do with this? Two things.

Our verdict

Rated 1,400W. Measured peak across 21 workloads: 1057.7W, on FLUX.1 Kontext at 99% utilisation. Measured LLM average: 286.2-391.6W at 12.2-40.3% utilisation, because single-stream token generation is bandwidth-bound and leaves Blackwell Ultra's compute idle. Peak temperature anywhere: 57°C. And tokens-per-watt collapses from 1.216 on GPT-OSS 20B to 0.127 on Llama 3.3 70B, roughly 10× the energy per token on the same silicon. Size your facility for the rating; size your expectations for the measurement. If your workload can't get a B300 past 400W, you don't need a B300.

FAQ

What is the NVIDIA B300's TDP?
1,400W is the board rating, a rating, not a measurement. It's the envelope the part is designed and cooled for. In our 1 Hz telemetry across 21 workloads, the highest instantaneous draw we recorded was 1057.7W, on FLUX.1 Kontext, which held the GPU at 99% utilisation. We never saw the chip approach its rating on any workload in our suite.
How much power does a B300 use running an LLM?
Far less than you'd expect. Measured averages: 286.2W on GPT-OSS 20B, 294.5W on Qwen3 4B, 338.3W on Llama 3.1 8B, 345.9W on Qwen3 32B, 352.5W on Gemma 3 27B, 367.2W on Qwen2.5-Coder 14B, 377.8W on Llama 3.3 70B, 391.6W on Phi-4 14B. A band of roughly 286.2-391.6W, between 20% and 28% of the card's 1,400W rating, across every language model we tested, from 4B to 70B.
Why does a B300 use so little power on language models?
Because single-stream token generation is bound by memory bandwidth, not compute. The GPU spends most of its time waiting on HBM reads rather than doing maths, so the tensor cores that account for most of the power budget sit largely idle. Our utilisation figures show it: 12.2% on GPT-OSS 20B, 31.3% on Llama 3.3 70B, 40.3% on Qwen3 14B. Compare that to FLUX.1 Kontext at 99% utilisation drawing 1029.5W. Same silicon, the workload decides everything.
What workload draws the most power on a B300?
Image generation and editing, by a wide margin. Our top five by average draw: FLUX.1 Kontext at 1029.5W (99% utilisation), Qwen-Image-Edit at 1019.6W (96.9%), Wan 2.2 video at 1012.6W (95.9%), FLUX.1-dev at 988.8W (98%), Qwen-Image at 944.6W (92.2%). These are diffusion workloads, dense tensor maths that genuinely saturates the chip. They're the only things in our suite that get a B300 anywhere near its rating.
How hot does a B300 get?
Not very, in our runs. Peak temperature across all 21 workloads was 57°C, on Wan 2.2 video. Language models ran between 39°C and 46°C. The hottest image workload, FLUX.1 Kontext, peaked at 56°C while pulling over a kilowatt. That's a testament to the cooling on the rented node rather than to the chip itself. But it does suggest that in a properly cooled environment, thermals are not the limiting factor. VRAM capacity and workload shape are.
Should I size my facility for 1,400W per B300 or for what you measured?
Size for the rating. Our numbers describe a single-GPU inference suite, and sustained multi-GPU FP4 training is a different animal, the rating exists because NVIDIA knows what the part can pull under conditions we didn't test. Reported GB300 NVL72 racks draw around 120kW, roughly consistent with the rating across 72 GPUs. What our data should change is your expectations, not your electrical plan: if your deployment is inference-shaped rather than training-shaped, expect to run far below your provisioned envelope, and be suspicious of any capacity model that assumes otherwise.
What is the B300's tokens-per-watt?
It falls off sharply as models grow. Measured: 1.216 tok/W on GPT-OSS 20B, 1.133 on Qwen3 4B, 0.851 on Llama 3.1 8B, 0.831 on DeepSeek-R1 Distill 8B, 0.46 on Qwen3 14B, 0.432 on Qwen2.5-Coder 14B, 0.343 on Mistral Small 24B, 0.265 on Gemma 3 27B, 0.242 on Qwen3 32B, and 0.127 on Llama 3.3 70B. A 70B costs roughly 10× as much energy per token as GPT-OSS 20B on the same chip. Larger models are not just slower, they're dramatically less efficient per token produced.

How we test

Every number on this page describing a single B300 is our own measurement. We rented a B300 and ran the GPU Battle AI Suite v2 across 21 workloads, the core 12 that every GPU on this site runs, plus 9 Blackwell Ultra extras. LLMs run on llama.cpp (llama-bench) at Q4_K_M with -p 512 -n 128; diffusion and video run on diffusers/ComfyUI at BF16 (SDXL at FP16). Each workload gets a warmup pass plus multiple timed runs (5 for small LLMs, 3 for large models and images, 2 for video). We publish the mean as the result and the minimum as the 1% low. Run-to-run variance across our fleet is under 0.5%. Telemetry is sampled at 1 Hz from nvidia-smi for the duration of every run: power, temperature, utilisation, clocks and peak VRAM. Every wattage, temperature and tokens-per-watt figure on this page is logged draw, not a board rating. The B300 was measured on 2026-07-12 on a CUDA 13 stack (harness 2.1.0-b300-cuda13, driver 580.95.05, torch 2.13.0+cu130), which differs from the CUDA 12.8 stack the rest of our fleet runs. We label it rather than hide it. The caveat that matters most: these are single-GPU, single-stream, batch-size-1 numbers. That is the honest way to measure what one chip does, and it is deliberately not how a datacenter runs a B300. Vendor and MLPerf figures use large batches across many GPUs and will be far higher. Neither is wrong, they answer different questions. Ours answers 'what does one of these actually do'. System-level specifications (GPU counts, memory totals, rack power) come from NVIDIA's published documentation and are cited as such; we have not taken a rack apart. Where NVIDIA's own published system figures disagree with each other, we show the arithmetic rather than pick a side.