Measured anomaly · both cards run first-party · Updated July 2026
We rented and measured both cards on the same 12 workloads. The B300, newer, more memory, higher power budget, won 11 of them, by margins from +4.7% to +80.5%. Then it lost Stable Diffusion XL by 57.9%: 14.6 it/s against the B200's 23.06. We're publishing it because it's what the data says, and because the telemetry we logged alongside it turns out to be more interesting than the result.

288GB HBM3e at 8 TB/s. Wins 11 of our 12 core workloads against the B200, by +4.7% to +80.5%. Then loses Stable Diffusion XL by 37%. We measured both.
Best for: Anyone who needs the memory ceiling gone, and who won't be running FP16 SDXL as their main job.

192GB HBM3e, also 8 TB/s, 1,000W. Loses to the B300 on 11 of 12 workloads but wins Stable Diffusion XL outright, 23.06 it/s vs 14.6, at 97% utilisation against the B300's 38.5%.
Best for: SDXL-heavy pipelines today, and anyone who wants Blackwell on a stack that's had longer to settle.
On paper this shouldn't happen. The B300 is Blackwell Ultra: 288GB of HBM3e against the B200's 192GB, the same 8 TB/s of bandwidth, a 1,400W board rating against 1,000W, and NVIDIA's claimed 1.5× dense FP4 and 2× attention performance. It should win everything. And it nearly does. We ran both cards through the identical 12-workload suite, same models, same steps, same quantisation, same harness, and the B300 came out ahead on eleven of them.
B300 vs B200, the same 12 workloads, both measured by us
Zero line, anything below it is a B300 loss
The B300's lead grows with the workload: +4.7% on Llama 3.1 8B, +58.4% on FLUX.1-dev, +80.5% on FLUX.1 Kontext. Then Stable Diffusion XL goes the other way, hard.
So we went to the telemetry, which is the whole reason we log it. And the explanation isn't in the throughput number at all. It's in the utilisation.
The same SDXL run, side by side, every logged metric
| Metric | B200 | B300 | What it tells us |
|---|---|---|---|
| Result | 23.06 it/s | 14.6 it/s | B200 wins by 58% |
| GPU utilisation | 97.0% | 38.5% | The B300 was idle most of the run |
| SM clock (avg) | 1965 MHz | 2032 MHz | B300 clocked HIGHER, not throttling |
| Power draw (avg) | 718.5 W | 488.4 W | B300 drew less because it did less |
| Peak temperature | 44°C | 41°C | Neither card was thermally limited |
| Peak VRAM (torch) | 12.3 GB | 12.4 GB | Identical, not a memory problem |
| Energy per run | 2804.0 J | 3010.4 J | The B300 used MORE energy for less work |
| Images per kWh | 3851.6 | 3587.6 | B200 is more efficient here too |
| torch / CUDA | 2.7.0+cu128 | 2.13.0+cu130 | The one thing that isn't matched |
Same model, same 30 steps, same precision, same harness, same operator, two days apart.
Which points at the one variable we couldn't match. At the time of testing, the B300 required a CUDA 13 stack, torch 2.13.0+cu130, while our entire fleet, the B200 included, runs torch 2.7.0+cu128 on CUDA 12.8. Everything else about the two runs was identical. We want to be careful about how far we push that, because the evidence isn't unanimous. Stable Diffusion 1.5, which we only ran on the B300, also showed low utilisation, 32%, and it's BF16, not FP16. That could mean the problem is broader than one precision path, or it could simply mean SD 1.5 is too small and too fast to saturate a chip like this at all (it finishes an image in 0.91 seconds). We can't distinguish those two explanations from the data we have. What we can say is that every BF16 diffusion workload the B300 ran at a meaningful size behaved exactly as expected.
B300 utilisation across every diffusion and video workload. SDXL is the outlier
FLUX.1 Kontext, FLUX.1-dev, Wan 2.2, Qwen-Image-Edit and Z-Image all pin the B300 at 95-99%, and the B300 wins every one of them against the B200 by +22.7% to +80.5%. SDXL alone sits at 38.5%.
There's a broader reason we're publishing a result that makes our own flagship card look bad. We have twenty other numbers on the B300 that say it's the fastest thing we've ever measured. The only reason to believe any of them is that we didn't quietly drop the one that doesn't fit the story. A benchmark suite that never produces an inconvenient result isn't a benchmark suite. It's marketing with error bars. And practically: if you're building an SDXL pipeline today, this matters to you. On our numbers the older, cheaper, lower-power card is 58% faster at that specific job. That's a real procurement fact, and it exists for maybe one software release.
The B300 beats the B200 on 11 of our 12 measured workloads, up to +80.5% on FLUX.1 Kontext. And loses Stable Diffusion XL by 57.9%, 14.6 it/s against 23.06. The telemetry rules out the obvious causes: the B300 ran a higher clock (2032 vs 1965 MHz), a cool 41°C, and used the same 12.4GB on a 288GB card. It just never got fed, 38.5% utilisation against the B200's 97.0%. Our leading suspect is the CUDA 13 stack the B300 alone runs on, and we'll test that directly on the next sweep. Until then: if SDXL is your main workload, the older card is measurably faster at it.
Every number describing a single B200 or B300 on this page is our own measurement. Both cards were rented and run through the GPU Battle AI Suite v2, the same 12 core workloads, same models, same settings, same harness. LLMs run on llama.cpp (llama-bench) at Q4_K_M with -p 512 -n 128; diffusion and video run on diffusers/ComfyUI at BF16, with SDXL at FP16. Each workload gets a warmup pass plus multiple timed runs (5 for small LLMs, 3 for large models and images, 2 for video); we publish the mean as the result and the minimum as the 1% low. Telemetry is sampled at 1 Hz from nvidia-smi for the whole run: power, temperature, utilisation, SM clocks and peak VRAM. Every wattage, clock and utilisation figure here is logged, not a board rating. The one asymmetry, and it matters on this page: the B200 was measured on 2026-07-10 on harness 2.0.0 with torch 2.7.0+cu128 (CUDA 12.8). The B300 was measured on 2026-07-12 on harness 2.1.0-b300-cuda13 with torch 2.13.0+cu130 (CUDA 13), because at the time of testing that was the stack the card required. Everything else about the two runs is identical. We label this on every B300 page rather than hide it, and on this page it is the central variable. All figures are single-GPU, single-stream, batch-size-1. That is the honest way to measure what one chip does, and it is not how a datacenter runs these cards. Vendor and MLPerf numbers use large batches across many GPUs and will be far higher. Neither is wrong; they answer different questions.