Pricing & access · B300 measured first-party · Updated July 2026
A single NVIDIA B300 costs roughly $53,000. An eight-GPU DGX B300 runs $400,000-500,000, and a full 72-GPU GB300 NVL72 rack is $3-4 million. Those are the numbers, and for almost everyone reading this they're the wrong numbers, because a B300 isn't a thing you buy. It's a thing you rent by the hour. We know, because that's how we benchmarked one across 21 workloads without spending more than pocket change.

288GB of HBM3e at 8 TB/s. We ran it through 21 workloads on 2026-07-12: 47.97 tok/s on Llama 3.3 70B, 9.71 it/s on FLUX.1-dev, 720p Wan video at 2.94 frames/s. The only card on our board of 102 where nothing in the core suite fails to fit.
Best for: Understanding what one B300 actually does before you spec, rent or buy anything built from them.
There are two honest answers to 'how much does a B300 cost', and the useful one isn't the number. The number: roughly $53,000 for a single B300. Around $400,000-500,000 for a DGX B300: eight of them plus dual Xeon 6776P CPUs, NVSwitch fabric, networking and storage in a 10U chassis. Roughly $3-4 million for a GB300 NVL72 rack: 72 GPUs, 36 Grace CPUs, 20.7TB of pooled HBM3e, about 120kW, liquid cooling mandatory. The useful answer: you're not buying one, and price isn't the constraint. The constraint is that a B300 is an SXM6 module rated at 1,400W. There is no slot for it in your machine, no connector on your power supply, and no cooler you can buy that will move a kilowatt of heat out of a tower case. It is not a card that happens to be expensive. It is a component of a datacenter that happens to be sold separately.
You rent it. We did. Our entire B300 campaign, 21 workloads, every model in our LLM ladder from Qwen3 4B up to the full Llama 3.3 70B, five image models, two video models, all with 1 Hz power and thermal telemetry, cost a few dollars of rented compute. Not $53,000. Not $400,000. Here's what those few dollars bought, measured rather than quoted.
Every workload we measured on a rented B300, 21 for the price of lunch
| Workload | Result | Unit | Peak VRAM | Fits? |
|---|---|---|---|---|
| Qwen3 4B | 333.34 | tok/s | 3.1 GB | Yes |
| Llama 3.1 8B | 287.23 | tok/s | 5.3 GB | Yes |
| Qwen2.5-Coder 14B | 158.49 | tok/s | 9.0 GB | Yes |
| Qwen3 32B | 83.68 | tok/s | 19.1 GB | Yes |
| Llama 3.3 70B | 47.97 | tok/s | 40.2 GB | Yes |
| Stable Diffusion XL | 14.6 | it/s | 16.5 GB | Yes |
| Z-Image Turbo | 5.29 | it/s | 26.1 GB | Yes |
| FLUX.1 dev | 9.71 | it/s | 37.0 GB | Yes |
| FLUX.1 Kontext dev | 4.82 | it/s | 35.8 GB | Yes |
| Qwen-Image-Edit | 4.07 | it/s | 60.5 GB | Yes |
| LTX-Video (distilled) | 31.87 | frames/s | 60.5 GB | Yes |
| Wan 2.2 5B (720p) | 2.94 | frames/s | 37.0 GB | Yes |
The core 12, plus 9 Blackwell Ultra extras not shown. Every single one fit in 288GB, no offload, no quantisation compromise, no gates. Across our board of 102 AI GPUs, the B300 is the only card we can say that about.
The argument against ownership, power draw vs the 1,400W you're paying for
1,400W rating, the capacity you paid $53,000 for
Across all 21 workloads the B300 never exceeded 1057.7W. On language models it averaged 286.2-391.6W at 12.2-40.3% utilisation. A B300 running one job at a time is an idle B300 you paid $53,000 for.
The economics only work when the card is saturated: batched, concurrent, many-user production inference, or sustained training. If that's your workload, you're probably not reading a page about B300 pricing; you're already talking to NVIDIA's partner network. If it isn't, the honest recommendation is the one we followed: rent it, run your actual job on it, and find out what it does before anyone quotes you anything. We don't publish an hourly rate here on purpose. Rates drift with supply and region constantly, and a number hardcoded into an evergreen page is a number that's wrong by next quarter. The live rate is on the provider's site, which is where it should be.
A B300 is about $53,000; a DGX B300 is $400,000-500,000; a GB300 NVL72 rack is $3-4 million. None of those are prices you're going to pay, because a 1,400W SXM6 module isn't a thing an individual buys. It's a datacenter component sold into systems. The real answer is that you rent it, and our own experience is the proof: 21 workloads measured on a rented B300 for the cost of lunch. What the money actually buys isn't speed. It's 288GB of HBM3e and the total absence of the VRAM ceiling that defines every other card on our board of 102.
Every number on this page describing a single B300 is our own measurement. We rented a B300 and ran the GPU Battle AI Suite v2 across 21 workloads, the core 12 that every GPU on this site runs, plus 9 Blackwell Ultra extras. LLMs run on llama.cpp (llama-bench) at Q4_K_M with -p 512 -n 128; diffusion and video run on diffusers/ComfyUI at BF16 (SDXL at FP16). Each workload gets a warmup pass plus multiple timed runs (5 for small LLMs, 3 for large models and images, 2 for video). We publish the mean as the result and the minimum as the 1% low. Run-to-run variance across our fleet is under 0.5%. Telemetry is sampled at 1 Hz from nvidia-smi for the duration of every run: power, temperature, utilisation, clocks and peak VRAM. Every wattage, temperature and tokens-per-watt figure on this page is logged draw, not a board rating. The B300 was measured on 2026-07-12 on a CUDA 13 stack (harness 2.1.0-b300-cuda13, driver 580.95.05, torch 2.13.0+cu130), which differs from the CUDA 12.8 stack the rest of our fleet runs. We label it rather than hide it. The caveat that matters most: these are single-GPU, single-stream, batch-size-1 numbers. That is the honest way to measure what one chip does, and it is deliberately not how a datacenter runs a B300. Vendor and MLPerf figures use large batches across many GPUs and will be far higher. Neither is wrong, they answer different questions. Ours answers 'what does one of these actually do'. System-level specifications (GPU counts, memory totals, rack power) come from NVIDIA's published documentation and are cited as such; we have not taken a rack apart. Where NVIDIA's own published system figures disagree with each other, we show the arithmetic rather than pick a side.