Pricing & access · B300 measured first-party · Updated July 2026

How Much Does an NVIDIA B300 Cost?

A single NVIDIA B300 costs roughly $53,000. An eight-GPU DGX B300 runs $400,000-500,000, and a full 72-GPU GB300 NVL72 rack is $3-4 million. Those are the numbers, and for almost everyone reading this they're the wrong numbers, because a B300 isn't a thing you buy. It's a thing you rent by the hour. We know, because that's how we benchmarked one across 21 workloads without spending more than pocket change.

The GPU we measured
NVIDIA B300

NVIDIA B300

288GB of HBM3e at 8 TB/s. We ran it through 21 workloads on 2026-07-12: 47.97 tok/s on Llama 3.3 70B, 9.71 it/s on FLUX.1-dev, 720p Wan video at 2.94 frames/s. The only card on our board of 102 where nothing in the core suite fails to fit.

Pros
  • 47.97 tok/s on Llama 3.3 70B (measured, single stream)
  • 288GB HBM3e, 12/12 core workloads fit, zero offload
  • 8 TB/s bandwidth, the highest on our board
  • AI Score 93.8/100, the highest score on our board
Cons
  • Roughly $53,000 per GPU; no retail channel exists
  • 1,400W rating, but we never measured above 1057.7W
  • Single-stream LLM work runs it at 12.2-40.3% utilisation
  • Measured on a CUDA 13 stack, unlike the rest of our fleet

Best for: Understanding what one B300 actually does before you spec, rent or buy anything built from them.

~$53k
One B300 GPU
no retail channel exists
$400-500k
DGX B300 (8 GPUs)
complete system
$3-4M
GB300 NVL72 rack
72 GPUs, ~120kW
a few $
What our 21-workload run cost
rented, by the hour

There are two honest answers to 'how much does a B300 cost', and the useful one isn't the number. The number: roughly $53,000 for a single B300. Around $400,000-500,000 for a DGX B300: eight of them plus dual Xeon 6776P CPUs, NVSwitch fabric, networking and storage in a 10U chassis. Roughly $3-4 million for a GB300 NVL72 rack: 72 GPUs, 36 Grace CPUs, 20.7TB of pooled HBM3e, about 120kW, liquid cooling mandatory. The useful answer: you're not buying one, and price isn't the constraint. The constraint is that a B300 is an SXM6 module rated at 1,400W. There is no slot for it in your machine, no connector on your power supply, and no cooler you can buy that will move a kilowatt of heat out of a tower case. It is not a card that happens to be expensive. It is a component of a datacenter that happens to be sold separately.

You rent it. We did. Our entire B300 campaign, 21 workloads, every model in our LLM ladder from Qwen3 4B up to the full Llama 3.3 70B, five image models, two video models, all with 1 Hz power and thermal telemetry, cost a few dollars of rented compute. Not $53,000. Not $400,000. Here's what those few dollars bought, measured rather than quoted.

Every workload we measured on a rented B300, 21 for the price of lunch

Qwen3 4B333.34
Llama 3.1 8B287.23
Qwen2.5-Coder 14B158.49
Qwen3 32B83.68
Llama 3.3 70B47.97
Stable Diffusion XL14.6
Z-Image Turbo5.29
FLUX.1 dev9.71
FLUX.1 Kontext dev4.82
Qwen-Image-Edit4.07
LTX-Video (distilled)31.87
Wan 2.2 5B (720p)2.94
WorkloadResultUnitPeak VRAMFits?
Qwen3 4B333.34tok/s3.1 GBYes
Llama 3.1 8B287.23tok/s5.3 GBYes
Qwen2.5-Coder 14B158.49tok/s9.0 GBYes
Qwen3 32B83.68tok/s19.1 GBYes
Llama 3.3 70B47.97tok/s40.2 GBYes
Stable Diffusion XL14.6it/s16.5 GBYes
Z-Image Turbo5.29it/s26.1 GBYes
FLUX.1 dev9.71it/s37.0 GBYes
FLUX.1 Kontext dev4.82it/s35.8 GBYes
Qwen-Image-Edit4.07it/s60.5 GBYes
LTX-Video (distilled)31.87frames/s60.5 GBYes
Wan 2.2 5B (720p)2.94frames/s37.0 GBYes

The core 12, plus 9 Blackwell Ultra extras not shown. Every single one fit in 288GB, no offload, no quantisation compromise, no gates. Across our board of 102 AI GPUs, the B300 is the only card we can say that about.

The argument against ownership, power draw vs the 1,400W you're paying for

GPT-OSS 20B
286.2 W
Qwen3 4B
294.5 W
Llama 3.1 8B
338.3 W
Qwen3 32B
345.9 W
Llama 3.3 70B
377.8 W
FLUX.1 dev
988.8 W
FLUX.1 Kontext dev
1029.5 W

1,400W rating, the capacity you paid $53,000 for

Across all 21 workloads the B300 never exceeded 1057.7W. On language models it averaged 286.2-391.6W at 12.2-40.3% utilisation. A B300 running one job at a time is an idle B300 you paid $53,000 for.

The economics only work when the card is saturated: batched, concurrent, many-user production inference, or sustained training. If that's your workload, you're probably not reading a page about B300 pricing; you're already talking to NVIDIA's partner network. If it isn't, the honest recommendation is the one we followed: rent it, run your actual job on it, and find out what it does before anyone quotes you anything. We don't publish an hourly rate here on purpose. Rates drift with supply and region constantly, and a number hardcoded into an evergreen page is a number that's wrong by next quarter. The live rate is on the provider's site, which is where it should be.

Our verdict

A B300 is about $53,000; a DGX B300 is $400,000-500,000; a GB300 NVL72 rack is $3-4 million. None of those are prices you're going to pay, because a 1,400W SXM6 module isn't a thing an individual buys. It's a datacenter component sold into systems. The real answer is that you rent it, and our own experience is the proof: 21 workloads measured on a rented B300 for the cost of lunch. What the money actually buys isn't speed. It's 288GB of HBM3e and the total absence of the VRAM ceiling that defines every other card on our board of 102.

FAQ

How much does one NVIDIA B300 GPU cost?
Roughly $53,000 per GPU based on industry reporting. There's no MSRP in the consumer sense and no retail channel: B300s are sold into systems, through NVIDIA's partner network, and to cloud providers. If you're pricing one as an individual or a small team, the purchase price is essentially academic.
How much does a DGX B300 system cost?
Approximately $400,000 to $500,000 for the complete 8-GPU system, including the eight B300s, dual Intel Xeon 6776P CPUs, NVSwitch fabric, networking, storage, and NVIDIA's software stack and support. A full GB300 NVL72 rack, 72 GPUs and 36 Grace CPUs, is reported at $3-4 million and requires liquid cooling and roughly 120kW.
Why can't I just buy a B300?
There's no consumer channel, and the card is physically useless outside a datacenter. A B300 is an SXM6 module rated at 1,400W. It doesn't go in a PC: no PCIe slot, no power connector you own, no cooling in a normal case that can move that much heat. It's sold as part of a system to buyers who have the facility to run it.
Is it cheaper to rent a B300 than buy one?
For any workload short of continuous, fully-utilised production, overwhelmingly yes. We measured a B300 across 21 workloads, every LLM in our ladder up to 70B, five image models, two video models, and the compute cost a few dollars. Buying the same silicon costs roughly $53,000. Ownership break-even requires keeping the card genuinely busy, and our telemetry suggests that's harder than it sounds: single-stream inference used under 30% of the B300's power budget.
What do you get for $53,000 that a $2,000 GPU can't do?
Capacity, mostly. 288GB of HBM3e at 8 TB/s means nothing in our 12-workload core suite failed to fit, the only card on our board of which that's true. Speed is a smaller part of the story than people expect: single-stream Llama 3.1 8B measured 287.23 tok/s, a real lead but not a 26× one, because token generation is bandwidth-bound and every modern card is bandwidth-bound too. You're buying the ability to hold enormous models and batch enormous concurrency, not a bigger number on a single chat session.
How much does it cost to rent a B300 per hour?
Rates move constantly with supply, provider, region and commitment, which is why we don't hardcode a number. It would be wrong within a quarter and we'd rather point you at the live rate. As a concrete data point: our full 21-workload run on a rented B300, including every LLM up to 70B, five image models and two video models with telemetry logged throughout, cost a few dollars in total.
Is the B300 worth it over an H200, or even a 5090?
Depends on the workload, and I say this having rented all three. For pure text generation a 5090 performs surprisingly close to the B300: the giant card's edge really shows on image and video, the VRAM- and tensor-heavy work. The H200 is the sweet middle ground: it holds everything in our suite including the 70B class at a far friendlier hourly rate. Rent the B300 when the job is heavy visual generation or you need the capacity; rent the H200 for big LLM work; and if it's text on models that fit 32GB, the 5090 in your own desktop is closer than you'd believe.

How we test

Every number on this page describing a single B300 is our own measurement. We rented a B300 and ran the GPU Battle AI Suite v2 across 21 workloads, the core 12 that every GPU on this site runs, plus 9 Blackwell Ultra extras. LLMs run on llama.cpp (llama-bench) at Q4_K_M with -p 512 -n 128; diffusion and video run on diffusers/ComfyUI at BF16 (SDXL at FP16). Each workload gets a warmup pass plus multiple timed runs (5 for small LLMs, 3 for large models and images, 2 for video). We publish the mean as the result and the minimum as the 1% low. Run-to-run variance across our fleet is under 0.5%. Telemetry is sampled at 1 Hz from nvidia-smi for the duration of every run: power, temperature, utilisation, clocks and peak VRAM. Every wattage, temperature and tokens-per-watt figure on this page is logged draw, not a board rating. The B300 was measured on 2026-07-12 on a CUDA 13 stack (harness 2.1.0-b300-cuda13, driver 580.95.05, torch 2.13.0+cu130), which differs from the CUDA 12.8 stack the rest of our fleet runs. We label it rather than hide it. The caveat that matters most: these are single-GPU, single-stream, batch-size-1 numbers. That is the honest way to measure what one chip does, and it is deliberately not how a datacenter runs a B300. Vendor and MLPerf figures use large batches across many GPUs and will be far higher. Neither is wrong, they answer different questions. Ours answers 'what does one of these actually do'. System-level specifications (GPU counts, memory totals, rack power) come from NVIDIA's published documentation and are cited as such; we have not taken a rack apart. Where NVIDIA's own published system figures disagree with each other, we show the arithmetic rather than pick a side.