Claims vs measurements · both cards run first-party · Updated July 2026
Is the NVIDIA B300 as Good as NVIDIA Says?
NVIDIA says Blackwell Ultra delivers 1.5× the dense FP4 performance and 2× the attention performance of Blackwell, and up to 50× the AI factory output of Hopper. We rented a B300 and a B200 and ran both through the identical 12-workload suite. The short answer: the claims are true, they're not misleading, and they almost certainly don't apply to what you're going to do with the card.
288GB at 8 TB/s, 1,400W. Against a B200 on our bench: +7.7% on Llama 3.3 70B, +80.5% on FLUX.1 Kontext. Same measurement, two completely different verdicts on the same card.
Pros
+80.5% over a B200 on FLUX.1 Kontext, beats NVIDIA's 1.5× claim
288GB vs the B200's 192GB, the real generational step
Nothing in our 12-workload suite fails to fit
AI Score 93.8/100
Cons
Only +7.7% over a B200 on Llama 3.3 70B
Identical 8 TB/s bandwidth to the B200. LLM inference can't improve
Loses SDXL to the B200 by 36.7%
Roughly $53,000 per GPU
Best for: Compute-bound work, diffusion, video, batched FP4. Not for making a chatbot faster than a B200 does.
192GB, also 8 TB/s, 1,000W, cheaper. On single-stream LLM inference our measurements put it within 4.7-7.7% of a B300, because the thing that limits both is identical.
Pros
Within 7.7% of a B300 on Llama 3.3 70B (measured)
Identical 8 TB/s memory bandwidth
All 12 core workloads fit in 192GB
Beats the B300 outright on Stable Diffusion XL
Cons
80.5% behind on FLUX.1 Kontext
192GB vs 288GB
Best for: LLM serving where the B300's premium buys you single digits.
1.5×
What NVIDIA claims vs Blackwell
dense FP4; plus 2× attention
+7.7%
What we measured on Llama 3.3 70B
single stream, vs a B200
+80.5%
What we measured on FLUX.1 Kontext
same two cards, same session
8 TB/s
Bandwidth, on BOTH cards
this is the whole explanation
NVIDIA's Blackwell Ultra claims are specific and, as far as we can tell, accurate. 1.5× the dense FP4 FLOPS of Blackwell. 2× the attention performance. Up to 50× the AI factory output of a Hopper-based platform. Up to 2.6× the training performance of the previous generation on large models like DeepSeek R1. We rented a B300 and a B200 and ran both through the same twelve workloads, same models, same settings, same harness, two days apart. Here's what we got.
B300 vs B200. Measured gain, split by what limits the workload
Llama 3.1 8B
4.7 % faster
Qwen2.5-Coder 14B
5 % faster
Qwen3 32B
6.5 % faster
Llama 3.3 70B
7.7 % faster
Z-Image Turbo
22.7 % faster
FLUX.1 dev
58.4 % faster
FLUX.1 Kontext dev
80.5 % faster
Qwen-Image-Edit
67.5 % faster
Wan 2.2 5B (720p)
56.4 % faster
LTX-Video (distilled)
18.9 % faster
NVIDIA's 1.5× claim = +50%
The first four are language models. The last six are diffusion and video. Same two cards, same session, and the gap between the two groups is not subtle.
NVIDIA's claim against our measurement, workload by workload
Workload
Bound by
NVIDIA implies
We measured
Claim holds?
Llama 3.1 8B
Memory bandwidth
+50% (1.5×)
+4.7%
No, nowhere close
Qwen2.5-Coder 14B
Memory bandwidth
+50% (1.5×)
+5.0%
No, nowhere close
Qwen3 32B
Memory bandwidth
+50% (1.5×)
+6.5%
No, nowhere close
Llama 3.3 70B
Memory bandwidth
+50% (1.5×)
+7.7%
No, nowhere close
Z-Image Turbo
Tensor compute
+50% (1.5×)
+22.7%
Partly
FLUX.1 dev
Tensor compute
+50% (1.5×)
+58.4%
Yes
FLUX.1 Kontext dev
Tensor compute
+50% (1.5×)
+80.5%
Yes
Qwen-Image-Edit
Tensor compute
+50% (1.5×)
+67.5%
Yes
Wan 2.2 5B (720p)
Tensor compute
+50% (1.5×)
+56.4%
Yes
LTX-Video (distilled)
Tensor compute
+50% (1.5×)
+18.9%
Partly
Stable Diffusion XL
Tensor compute
+50% (1.5×)
-36.7%
No, B300 loses
NVIDIA's 1.5× figure is a dense FP4 tensor claim, not a promise about every workload. We map it across our suite to show where it lands and where it doesn't. The SDXL row is an anomaly we've written up separately, the telemetry suggests a software problem, not silicon.
So: is the B300 as good as NVIDIA says? Yes. And that's the wrong question. NVIDIA's claims are about tensor compute, measured on batched FP4 workloads across many GPUs. Nothing in our data contradicts them. On the workloads that are actually limited by tensor compute, the B300 hits or exceeds the 1.5× figure, FLUX.1 Kontext is 80% faster than a B200 on our bench, which is beyond what NVIDIA advertises. The question worth asking is whether the claim applies to you. If you're buying a B300 to serve a language model faster than a B200 does, our measurements say you're buying between five and eight percent. Not fifty. The bandwidth is identical, the workload is bandwidth-bound, and the FP4 improvements sit idle.
Why the LLM gain is small, B300 utilisation, by workload
GPT-OSS 20B
12.2 % util
Llama 3.3 70B
31.3 % util
Qwen3 14B
40.3 % util
Z-Image Turbo
95 % util
Qwen-Image-Edit
96.9 % util
FLUX.1 dev
98 % util
FLUX.1 Kontext dev
99 % util
Logged at 1 Hz on our bench. Running a 70B, the B300 sits at 31.3% utilisation, it spends most of its life waiting on memory, not computing. FLUX.1 Kontext pins it at 99.0%. You cannot benefit from compute you never use.
Our verdict
NVIDIA claims 1.5× dense FP4 and 2× attention for Blackwell Ultra, and nothing we measured contradicts that. On compute-bound work the B300 delivers: +58.4% over a B200 on FLUX.1-dev and +80.5% on FLUX.1 Kontext, which is beyond the headline claim. On LLM inference it delivers +4.7% to +7.7%: because both cards read memory at exactly 8 TB/s and token generation is bandwidth-bound, so the extra tensor throughput never gets used. Our utilisation logs show the B300 sitting at 31.3% while running a 70B. The claim is true. Whether it's true for you depends entirely on whether your workload is limited by compute or by memory: and if you're serving language models, it's the latter, and you're buying capacity rather than speed.
FAQ
Does the B300 really deliver 1.5× the performance of a B200?
On compute-bound workloads, yes and more, we measured +58.4% on FLUX.1-dev, +67.5% on Qwen-Image-Edit and +80.5% on FLUX.1 Kontext, all against a B200 on the identical test. On memory-bound workloads, no: +4.7% on Llama 3.1 8B and +7.7% on Llama 3.3 70B. NVIDIA's 1.5× figure is a dense FP4 tensor claim, and it's accurate, it just only applies where tensor compute is the limit.
Why is the B300 only 7.7% faster than a B200 at running Llama 70B?
Because both cards have exactly the same memory bandwidth: 8 TB/s. LLM token generation is bound by how fast the card can read the model out of VRAM, once per token, not by how fast it can do maths. Two cards that read memory at the same speed produce tokens at nearly the same speed, regardless of what their tensor cores can do. Our utilisation telemetry confirms it: running a 70B, the B300 sits at 31.3% utilisation. Most of the chip isn't doing anything, so making that part faster doesn't help.
Is NVIDIA's B300 marketing misleading?
We don't think so, and we'd say if we did. The claims are specific, 1.5× dense FP4, 2× attention, and they're about tensor compute measured on batched workloads. They're accurate on that basis, and our compute-bound measurements back them up. What happens is that buyers read a compute claim as a general performance claim. That's a reading error rather than a marketing lie, but it's an expensive one at roughly $53,000 a GPU, which is why we think publishing the split matters.
Should I buy a B300 or a B200 for LLM inference?
On our single-stream measurements the B300 is +7.7% on Llama 3.3 70B. That's what the premium buys you on speed. The genuine reason to choose a B300 is capacity: 288GB against 192GB, which means models and context lengths a B200 physically cannot hold, plus more headroom for batching. If your model fits comfortably in 192GB and you're not batching heavily, the B200 does nearly the same job. Buy the B300 for the memory, not the tok/s.
What about NVIDIA's 50× AI factory claim?
That's a rack-scale, batched, FP4, disaggregated-serving figure against a Hopper-based platform, a completely different measurement from anything we do. We test one GPU, one stream, batch size 1, because that's the honest way to say what a single chip does for a single user. We can't verify a rack claim and we don't try. Both numbers can be true; they answer different questions, and ours is the one that tells you what the card in front of you will do.
How did you test this?
We rented a B200 and a B300 and ran both through the identical GPU Battle AI Suite v2, same 12 workloads, same models, same settings, same harness. LLMs at Q4_K_M on llama.cpp, diffusion at BF16 on diffusers. Warmup plus multiple timed runs each, with 1 Hz telemetry logging power, temperature, utilisation and clocks throughout. Run-to-run variance was under 0.5%. One asymmetry we flag everywhere: the B300 required a CUDA 13 stack at test time while the B200 ran CUDA 12.8. That matters for the SDXL result specifically, which we've written up on its own page.
Is the B300 worth it over an H200, or even a 5090?
Depends on the workload, and I say this having rented all three. For pure text generation a 5090 performs surprisingly close to the B300: the giant card's edge really shows on image and video, the VRAM- and tensor-heavy work. The H200 is the sweet middle ground: it holds everything in our suite including the 70B class at a far friendlier hourly rate. Rent the B300 when the job is heavy visual generation or you need the capacity; rent the H200 for big LLM work; and if it's text on models that fit 32GB, the 5090 in your own desktop is closer than you'd believe.
How we test
We rented both a B200 and a B300 and ran them through the identical GPU Battle AI Suite v2, same 12 workloads, same models, same settings, same harness, two days apart. LLMs run on llama.cpp (llama-bench) at Q4_K_M with -p 512 -n 128; diffusion and video run on diffusers/ComfyUI at BF16, with SDXL at FP16. Each workload gets a warmup pass plus multiple timed runs; we publish the mean. Run-to-run variance is under 0.5%. Telemetry, power, temperature, utilisation, SM clocks, peak VRAM, is sampled at 1 Hz. NVIDIA's claims quoted here come from NVIDIA's own published material and are cited as claims, not verified by us at rack scale. We have not tested a DGX or an NVL72; we test single GPUs. The one asymmetry between the two runs: the B200 was measured on 2026-07-10 with torch 2.7.0+cu128 (CUDA 12.8), the B300 on 2026-07-12 with torch 2.13.0+cu130 (CUDA 13), because that was the stack the card required. Everything else matched. This is material to the Stable Diffusion XL result and we treat it separately. All figures are single-GPU, single-stream, batch-size-1. NVIDIA's performance claims are batched, multi-GPU and FP4. Neither measurement is wrong and neither invalidates the other, they answer different questions. Ours answers what one B300 does for one user. NVIDIA's answers what a fleet of them does for a datacenter. The point of this page is that those two answers are much further apart than most buyers assume.