Measured anomaly · both cards run first-party · Updated July 2026

Why Is the B300 Slower Than the B200 at Stable Diffusion?

We rented and measured both cards on the same 12 workloads. The B300, newer, more memory, higher power budget, won 11 of them, by margins from +4.7% to +80.5%. Then it lost Stable Diffusion XL by 57.9%: 14.6 it/s against the B200's 23.06. We're publishing it because it's what the data says, and because the telemetry we logged alongside it turns out to be more interesting than the result.

Blackwell Ultra
NVIDIA B300

NVIDIA B300

288GB HBM3e at 8 TB/s. Wins 11 of our 12 core workloads against the B200, by +4.7% to +80.5%. Then loses Stable Diffusion XL by 37%. We measured both.

Pros
  • 47.97 tok/s on Llama 3.3 70B (measured)
  • 9.71 it/s on FLUX.1-dev, 58% faster than a B200
  • 288GB, nothing in our suite fails to fit
  • AI Score 93.8/100
Cons
  • Loses SDXL to the B200: 14.6 vs 23.06 it/s
  • Only 38.5% utilised on that run, the chip sat idle
  • Roughly $53,000 per GPU
  • Measured on a CUDA 13 stack, unlike our CUDA 12.8 fleet

Best for: Anyone who needs the memory ceiling gone, and who won't be running FP16 SDXL as their main job.

Blackwell
NVIDIA B200

NVIDIA B200

192GB HBM3e, also 8 TB/s, 1,000W. Loses to the B300 on 11 of 12 workloads but wins Stable Diffusion XL outright, 23.06 it/s vs 14.6, at 97% utilisation against the B300's 38.5%.

Pros
  • 23.06 it/s on SDXL, beats the B300 by 58%
  • 44.54 tok/s on Llama 3.3 70B (measured)
  • All 12 core workloads fit in 192GB
  • AI Score 78.0/100 · mature CUDA 12.8 stack
Cons
  • Slower than the B300 on the other 11 workloads
  • 192GB vs 288GB
  • 1,000W board rating

Best for: SDXL-heavy pipelines today, and anyone who wants Blackwell on a stack that's had longer to settle.

11 of 12
Workloads the B300 wins
by +4.7% to +80.5%
57.9%
How much it loses SDXL by
14.6 vs 23.06 it/s
38.5%
B300 GPU utilisation on SDXL
the B200 hit 97.0%
2032MHz
B300 SM clock on that run
HIGHER than the B200's 1965

On paper this shouldn't happen. The B300 is Blackwell Ultra: 288GB of HBM3e against the B200's 192GB, the same 8 TB/s of bandwidth, a 1,400W board rating against 1,000W, and NVIDIA's claimed 1.5× dense FP4 and 2× attention performance. It should win everything. And it nearly does. We ran both cards through the identical 12-workload suite, same models, same steps, same quantisation, same harness, and the B300 came out ahead on eleven of them.

B300 vs B200, the same 12 workloads, both measured by us

Qwen3 4B
4.8 % faster
Llama 3.1 8B
4.7 % faster
Qwen2.5-Coder 14B
5 % faster
Qwen3 32B
6.5 % faster
Llama 3.3 70B
7.7 % faster
Stable Diffusion XL
-36.7 % faster
Z-Image Turbo
22.7 % faster
FLUX.1 dev
58.4 % faster
FLUX.1 Kontext dev
80.5 % faster
Qwen-Image-Edit
67.5 % faster
LTX-Video (distilled)
18.9 % faster
Wan 2.2 5B (720p)
56.4 % faster

Zero line, anything below it is a B300 loss

The B300's lead grows with the workload: +4.7% on Llama 3.1 8B, +58.4% on FLUX.1-dev, +80.5% on FLUX.1 Kontext. Then Stable Diffusion XL goes the other way, hard.

So we went to the telemetry, which is the whole reason we log it. And the explanation isn't in the throughput number at all. It's in the utilisation.

The same SDXL run, side by side, every logged metric

Result23.06 it/s
GPU utilisation97.0%
SM clock (avg)1965 MHz
Power draw (avg)718.5 W
Peak temperature44°C
Peak VRAM (torch)12.3 GB
Energy per run2804.0 J
Images per kWh3851.6
torch / CUDA2.7.0+cu128
MetricB200B300What it tells us
Result23.06 it/s14.6 it/sB200 wins by 58%
GPU utilisation97.0%38.5%The B300 was idle most of the run
SM clock (avg)1965 MHz2032 MHzB300 clocked HIGHER, not throttling
Power draw (avg)718.5 W488.4 WB300 drew less because it did less
Peak temperature44°C41°CNeither card was thermally limited
Peak VRAM (torch)12.3 GB12.4 GBIdentical, not a memory problem
Energy per run2804.0 J3010.4 JThe B300 used MORE energy for less work
Images per kWh3851.63587.6B200 is more efficient here too
torch / CUDA2.7.0+cu1282.13.0+cu130The one thing that isn't matched

Same model, same 30 steps, same precision, same harness, same operator, two days apart.

Which points at the one variable we couldn't match. At the time of testing, the B300 required a CUDA 13 stack, torch 2.13.0+cu130, while our entire fleet, the B200 included, runs torch 2.7.0+cu128 on CUDA 12.8. Everything else about the two runs was identical. We want to be careful about how far we push that, because the evidence isn't unanimous. Stable Diffusion 1.5, which we only ran on the B300, also showed low utilisation, 32%, and it's BF16, not FP16. That could mean the problem is broader than one precision path, or it could simply mean SD 1.5 is too small and too fast to saturate a chip like this at all (it finishes an image in 0.91 seconds). We can't distinguish those two explanations from the data we have. What we can say is that every BF16 diffusion workload the B300 ran at a meaningful size behaved exactly as expected.

B300 utilisation across every diffusion and video workload. SDXL is the outlier

Stable Diffusion XL
38.5 % util
LTX-Video (distilled)
65.8 % util
Z-Image Turbo
95 % util
Qwen-Image-Edit
96.9 % util
FLUX.1 dev
98 % util
Wan 2.2 5B (720p)
95.9 % util
FLUX.1 Kontext dev
99 % util

FLUX.1 Kontext, FLUX.1-dev, Wan 2.2, Qwen-Image-Edit and Z-Image all pin the B300 at 95-99%, and the B300 wins every one of them against the B200 by +22.7% to +80.5%. SDXL alone sits at 38.5%.

There's a broader reason we're publishing a result that makes our own flagship card look bad. We have twenty other numbers on the B300 that say it's the fastest thing we've ever measured. The only reason to believe any of them is that we didn't quietly drop the one that doesn't fit the story. A benchmark suite that never produces an inconvenient result isn't a benchmark suite. It's marketing with error bars. And practically: if you're building an SDXL pipeline today, this matters to you. On our numbers the older, cheaper, lower-power card is 58% faster at that specific job. That's a real procurement fact, and it exists for maybe one software release.

Our verdict

The B300 beats the B200 on 11 of our 12 measured workloads, up to +80.5% on FLUX.1 Kontext. And loses Stable Diffusion XL by 57.9%, 14.6 it/s against 23.06. The telemetry rules out the obvious causes: the B300 ran a higher clock (2032 vs 1965 MHz), a cool 41°C, and used the same 12.4GB on a 288GB card. It just never got fed, 38.5% utilisation against the B200's 97.0%. Our leading suspect is the CUDA 13 stack the B300 alone runs on, and we'll test that directly on the next sweep. Until then: if SDXL is your main workload, the older card is measurably faster at it.

FAQ

Is the B300 slower than the B200?
On one workload out of twelve, yes, measurably. Stable Diffusion XL runs at 14.6 it/s on the B300 versus 23.06 it/s on the B200, so the B200 is 57.9% faster. On the other eleven the B300 wins, from +4.7% on Llama 3.1 8B up to +80.5% on FLUX.1 Kontext. Both sets of numbers are our own measurements on the same suite.
Why is the B300 slower at SDXL?
We don't know for certain, and we'd rather say that than guess confidently. What our telemetry rules out is the hardware: on that run the B300 held a higher SM clock than the B200 (2032 vs 1965 MHz), peaked at only 41°C, and used the same ~12.4GB of VRAM on a card with 288GB. It simply wasn't busy, 38.5% average utilisation against the B200's 97.0%. That pattern says the GPU wasn't being fed work fast enough, which is a software characteristic, not a silicon one. Our leading suspect is the CUDA 13 stack the B300 required at test time, versus the CUDA 12.8 stack the B200 and the rest of our fleet run.
Could this just be a bad run or a measurement error?
We don't think so. We ran SDXL three times on each card. Standard deviation was 0.014 on the B200 and 0.044 on the B300, the three B300 runs came in at 14.548, 14.656 and 14.602 it/s. The result is extremely repeatable. Everything else about the two runs matched: same model, same 30 steps, same FP16 precision, same harness, same driver version. The reproducibility is exactly why we treat it as a finding rather than an outlier to discard.
Should I buy a B200 instead of a B300 for Stable Diffusion?
If SDXL specifically is your production workload, our numbers say the B200 does it 57.9% faster, at lower power, using less energy per image (2804.0 J versus 3010.4 J), so yes, on today's software. Two caveats. First, if this is a stack problem it may be fixed by a software update, and then the advantage evaporates. Second, SDXL is one workload: on FLUX.1-dev the B300 is 58% faster, on FLUX.1 Kontext 81% faster. If your pipeline is modern diffusion rather than SDXL, the B300 wins comfortably.
Does the B300 beat the B200 on FLUX?
Decisively. FLUX.1-dev: 9.71 it/s versus 6.13. The B300 is 58% faster. FLUX.1 Kontext: 4.82 versus 2.67, 81% faster. Qwen-Image-Edit: 4.07 versus 2.43, 67% faster. All of those run BF16, and on all of them the B300 sits at 95-99% utilisation. The chip is working. Only SDXL, our lone FP16 workload, shows the collapse.
Why publish a result that makes your own numbers look inconsistent?
Because the alternative is worse. We have twenty other B300 figures saying it's the fastest card we've ever measured, and the only reason anyone should believe them is that we didn't hide the one that disagrees. A suite that never produces an inconvenient result isn't measuring anything. It's also useful: if you're running SDXL in production right now, a 58% gap in favour of the older card is a real fact about your procurement, whatever its cause turns out to be.

How we test

Every number describing a single B200 or B300 on this page is our own measurement. Both cards were rented and run through the GPU Battle AI Suite v2, the same 12 core workloads, same models, same settings, same harness. LLMs run on llama.cpp (llama-bench) at Q4_K_M with -p 512 -n 128; diffusion and video run on diffusers/ComfyUI at BF16, with SDXL at FP16. Each workload gets a warmup pass plus multiple timed runs (5 for small LLMs, 3 for large models and images, 2 for video); we publish the mean as the result and the minimum as the 1% low. Telemetry is sampled at 1 Hz from nvidia-smi for the whole run: power, temperature, utilisation, SM clocks and peak VRAM. Every wattage, clock and utilisation figure here is logged, not a board rating. The one asymmetry, and it matters on this page: the B200 was measured on 2026-07-10 on harness 2.0.0 with torch 2.7.0+cu128 (CUDA 12.8). The B300 was measured on 2026-07-12 on harness 2.1.0-b300-cuda13 with torch 2.13.0+cu130 (CUDA 13), because at the time of testing that was the stack the card required. Everything else about the two runs is identical. We label this on every B300 page rather than hide it, and on this page it is the central variable. All figures are single-GPU, single-stream, batch-size-1. That is the honest way to measure what one chip does, and it is not how a datacenter runs these cards. Vendor and MLPerf numbers use large batches across many GPUs and will be far higher. Neither is wrong; they answer different questions.