NVIDIA B300, AI & Machine Learning Benchmarks & Specs

288GB · AI Score 93.8/100 · first-party measured on 12 AI workloads

93.8 AI Score ✓ Measured

Every number on this page is first-party: NVIDIA B300 was run on our pinned 12-workload AI suite on 2026-07-12, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) NVIDIA B300 delivers about 287.23 tokens/sec. Stepping up to Qwen3 32B it holds roughly 83.68 tok/s. The full Llama 3.3 70B still runs, at about 47.97 tok/s. For image generation, SDXL runs at 14.6 it/s, and FLUX.1-dev at 9.71 it/s. All 12 workloads fit in 288GB. There is no model in our suite this card has to turn down. NVIDIA B300 isn't a retail purchase for most people. It's rented by the hour. You can run this exact card on RunPod.

AI & Machine Learning benchmark results

Text Generation tok/s 123

LFM2.5-1.2B955.13
gemma-3-270m926.67
Qwen1.5-0.5B894.84
Llama 3.2 1B889.19
SmolLM2-135M855.87
Qwen2.5-Coder-0.5B843.44
Qwen2.5-0.5B843.4
Qwen3-0.6B742.64
Qwen3 0.6B718.58
LFM2.5-8B-A1B596.16
Qwen2.5-1.5B584.96
Qwen2.5-Coder-1.5B584.5
WorkloadResultTelemetryData
LFM2.5-1.2B955.13 tok/s
259 W38°CQ4_K_M
✓ Measured
gemma-3-270m926.67 tok/s
242 W33°CQ4_K_M
✓ Measured
Qwen1.5-0.5B894.84 tok/s
246 W34°CQ4_K_M
✓ Measured
Llama 3.2 1B889.19 tok/s
263 W44°CQ4_K_M
✓ Measured
SmolLM2-135M855.87 tok/s
244 W34°CQ4_K_M
✓ Measured
Qwen2.5-Coder-0.5B843.44 tok/s
244 W33°CQ4_K_M
✓ Measured
Qwen2.5-0.5B843.4 tok/s
252 W33°CQ4_K_M
✓ Measured
Qwen3-0.6B742.64 tok/s
249 W34°CQ4_K_M
✓ Measured
Qwen3 0.6B718.58 tok/s
223 W42°CQ4_K_M
✓ Measured
LFM2.5-8B-A1B596.16 tok/s
277 W36°CQ4_K_M
✓ Measured
Qwen2.5-1.5B584.96 tok/s
267 W35°CQ4_K_M
✓ Measured
Qwen2.5-Coder-1.5B584.5 tok/s
265 W37°CQ4_K_M
✓ Measured
Qwen2-1.5B583.79 tok/s
265 W36°CQ4_K_M
✓ Measured
Qwen3-1.7B575.97 tok/s
274 W37°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 1.5B562.14 tok/s
274 W44°CQ4_K_M
✓ Measured
Qwen3 1.7B555.1 tok/s
275 W44°CQ4_K_M
✓ Measured
gemma-3-1b533.32 tok/s
271 W35°CQ4_K_M
✓ Measured
Llama-3.2-3B-Instruct-uncensored448.94 tok/s
323 W39°CQ4_K_M
✓ Measured
Hermes-3-Llama-3.2-3B446.78 tok/s
257 W37°CQ4_K_M
✓ Measured
Llama 3.2 3B432.24 tok/s
268 W46°CQ4_K_M
✓ Measured
SmolLM3-3B420.3 tok/s
321 W37°CQ4_K_M
✓ Measured
gemma-2-2b-it-abliterated418.18 tok/s
305 W37°CQ4_K_M
✓ Measured
gemma-2-2b417.84 tok/s
307 W36°CQ4_K_M
✓ Measured
SmolLM3 3B411.25 tok/s
278 W46°CQ4_K_M
✓ Measured
Qwen2.5-Coder-3B410.06 tok/s
317 W38°CQ4_K_M
✓ Measured
Phi-4-mini409.58 tok/s
324 W39°CQ4_K_M
✓ Measured
Qwen2.5-3B409.56 tok/s
300 W37°CQ4_K_M
✓ Measured
Phi-4 Mini 3.8B398.16 tok/s
301 W47°CQ4_K_M
✓ Measured
AI21-Jamba-Reasoning-3B380.88 tok/s
311 W37°CQ4_K_M
✓ Measured
phi-2356.65 tok/s
316 W37°CQ4_K_M
✓ Measured
gpt-oss-20b355.56 tok/s
279 W37°CQ4_K_M
✓ Measured
Phi-3.5-mini352.25 tok/s
334 W37°CQ4_K_M
✓ Measured
Nemotron-3-Nano-30B-A3B345.31 tok/s
280 W37°CQ4_K_M
✓ Measured
DeepSeek-Coder-V2-Lite340.53 tok/s
279 W37°CQ4_K_M
✓ Measured
Qwen3-4B-Instruct-2507339.78 tok/s
325 W38°CQ4_K_M
✓ Measured
Qwen3-4B-Thinking-2507339.74 tok/s
325 W37°CQ4_K_M
✓ Measured
Qwen3-4B-Instruct-2507339.72 tok/s
313 W38°CQ4_K_M
✓ Measured
Qwen3 4B333.34 tok/s
3.1 GB peak295 W39°C1.13 tok/WQ4_K_M
✓ Measured
Llama-2-7B315.9 tok/s
379 W40°CQ4_K_M
✓ Measured
Mistral-7B-Instruct-v0.2306.79 tok/s
377 W40°CQ4_K_M
✓ Measured
Mistral-7B-Instruct-v0.3306.46 tok/s
366 W40°CQ4_K_M
✓ Measured
Mistral-7B-Instruct-v0.1306.44 tok/s
372 W40°CQ4_K_M
✓ Measured
Gemma 3 4B301.07 tok/s
283 W46°CQ4_K_M
✓ Measured
Mistral 7B v0.3298.95 tok/s
313 W48°CQ4_K_M
✓ Measured
Qwen2.5-7B293.63 tok/s
370 W39°CQ4_K_M
✓ Measured
L3-8B-Stheno-v3.2293 tok/s
381 W40°CQ4_K_M
✓ Measured
Meta-Llama-3.1-8B293 tok/s
352 W40°CQ4_K_M
✓ Measured
DarkIdol-Llama-3.1-8B-Instruct-1.2-Uncensored292.55 tok/s
361 W41°CQ4_K_M
✓ Measured
Qwen2.5-Coder-7B-Instruct-abliterated292.21 tok/s
370 W39°CQ4_K_M
✓ Measured
Qwen3-30B-A3B291.82 tok/s
287 W36°CQ4_K_M
✓ Measured
dolphin-2.9-llama3-8b290.47 tok/s
354 W40°CQ4_K_M
✓ Measured
Llama 3.1 8B287.23 tok/s
5.3 GB peak338 W41°C0.85 tok/WQ4_K_M
✓ Measured
Dolphin X1 8B287.02 tok/s
292 W40°CQ4_K_M
✓ Measured
Qwen2.5-Coder 7B286.5 tok/s
299 W48°CQ4_K_M
✓ Measured
Dolphin 3.0 Llama 3.1 8B285.88 tok/s
300 W40°CQ4_K_M
✓ Measured
Qwen3-Coder 30B A3B284.88 tok/s
256 W36°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 7B284.77 tok/s
291 W42°CQ4_K_M
✓ Measured
Qwen3 30B A3B284.66 tok/s
250 W39°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill Llama 8B283.77 tok/s
306 W42°CQ4_K_M
✓ Measured
DeepSeek-R1-0528-Qwen3-8B270.35 tok/s
375 W40°CQ4_K_M
✓ Measured
Qwen3-8B270.35 tok/s
383 W41°CQ4_K_M
✓ Measured
Josiefied-Qwen3-8B-abliterated-v1270.09 tok/s
362 W40°CQ4_K_M
✓ Measured
Qwen3 30B A3B (Q3_K_M)268.71 tok/s
302 W37°CQ3_K_M
✓ Measured
Qwen3 8B261.4 tok/s
275 W41°CQ4_K_M
✓ Measured
Dolphin X1 Trinity Nano 6B257.93 tok/s
249 W33°CQ4_K_M
✓ Measured
Ornith-1.0-9B240.67 tok/s
377 W40°CQ4_K_M
✓ Measured
KAT-Coder-V2.5-Dev239.66 tok/s
287 W36°CQ4_K_M
✓ Measured
Qwen-AgentWorld-35B-A3B233.04 tok/s
293 W36°CQ4_K_M
✓ Measured
Ornith-1.0-35B232.36 tok/s
291 W36°CQ4_K_M
✓ Measured
gemma-2-9b202.22 tok/s
386 W40°CQ4_K_M
✓ Measured
Qwen3-Next-80B-A3B-Thinking200.55 tok/s
287 W36°CQ4_K_M
✓ Measured
Qwen3-Coder-Next199.39 tok/s
288 W36°CQ4_K_M
✓ Measured
Qwen3-Coder-Next-abliterated199.12 tok/s
280 W36°CQ4_K_M
✓ Measured
NemoMix-Unleashed-12B197.06 tok/s
403 W41°CQ4_K_M
✓ Measured
Mistral-Nemo-Instruct-2407196.92 tok/s
406 W41°CQ4_K_M
✓ Measured
GLM-4.7-Flash196.52 tok/s
303 W35°CQ4_K_M
✓ Measured
Qwen3-Coder-Next196.34 tok/s
292 W36°CQ4_K_M
✓ Measured
Qwen3-Next-80B-A3B-Thinking194.98 tok/s
292 W36°CQ4_K_M
✓ Measured
Qwen3-Next-80B-A3B191.32 tok/s
284 W36°CQ4_K_M
✓ Measured
GLM-4.7-Flash-REAP-23B-A3B178.49 tok/s
300 W36°CQ4_K_M
✓ Measured
Phi-4 14B177.41 tok/s
334 W44°CQ4_K_M
✓ Measured
Qwen3-14B169.52 tok/s
422 W41°CQ4_K_M
✓ Measured
Hermes-4-14B169.51 tok/s
422 W41°CQ4_K_M
✓ Measured
Qwen3 14B165.44 tok/s
315 W43°CQ4_K_M
✓ Measured
Phi-4 14B (Q3_K_M)165.06 tok/s
439 W42°CQ3_K_M
✓ Measured
Gemma 3 12B161.27 tok/s
315 W42°CQ4_K_M
✓ Measured
Qwen2.5-Coder-14B-Instruct-abliterated159.5 tok/s
423 W40°CQ4_K_M
✓ Measured
Uncensored159.48 tok/s
416 W41°CQ4_K_M
✓ Measured
EVA-Qwen2.5-14B-v0.2159.46 tok/s
416 W41°CQ4_K_M
✓ Measured
Qwen2.5-14B159.33 tok/s
424 W40°CQ4_K_M
✓ Measured
Qwen2.5-Coder 14B158.49 tok/s
9 GB peak367 W42°C0.43 tok/WQ4_K_M
✓ Measured
DeepSeek-R1 Distill 14B156.33 tok/s
328 W43°CQ4_K_M
✓ Measured
Gemma 4 12B155.96 tok/s
314 W42°CQ4_K_M
✓ Measured
Gemma 3 12B (Q3_K_M)149.61 tok/s
412 W40°CQ3_K_M
✓ Measured
DeepSeek-R1 Distill 14B (Q3_K_M)144.74 tok/s
424 W41°CQ3_K_M
✓ Measured
StarCoder2 15B144.05 tok/s
330 W41°CQ4_K_M
✓ Measured
Cydonia-24B-v4.3121.9 tok/s
477 W43°CQ4_K_M
✓ Measured
Dolphin-Mistral-24B-Venice-Edition121.87 tok/s
473 W43°CQ4_K_M
✓ Measured
Devstral Small 24B121.18 tok/s
323 W42°CQ4_K_M
✓ Measured
Dolphin 3.0 R1 Mistral 24B120.95 tok/s
348 W43°CQ4_K_M
✓ Measured
Dolphin Mistral 24B Venice120.93 tok/s
331 W43°CQ4_K_M
✓ Measured
Codestral 22B119.81 tok/s
328 W42°CQ4_K_M
✓ Measured
Mistral Small 24B119.3 tok/s
345 W45°CQ4_K_M
✓ Measured
Mistral Small 24B (Q3_K_M)108.37 tok/s
460 W43°CQ3_K_M
✓ Measured
Codestral 22B (Q3_K_M)107.98 tok/s
465 W43°CQ3_K_M
✓ Measured
Gemma 3 27B93.44 tok/s
326 W42°CQ4_K_M
✓ Measured
Olmo-3.1-32B-Think87.85 tok/s
476 W43°CQ4_K_M
✓ Measured
Dolphin 2.9.1 Yi 1.5 34B85.15 tok/s
312 W43°CQ4_K_M
✓ Measured
Qwen3 32B83.68 tok/s
19.1 GB peak346 W44°C0.24 tok/WQ4_K_M
✓ Measured
Qwen2.5-32B83.41 tok/s
490 W43°CQ4_K_M
✓ Measured
DeepSeek-R1-Distill-Qwen-32B-abliterated83.19 tok/s
468 W43°CQ4_K_M
✓ Measured
QwQ 32B82.91 tok/s
328 W43°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 32B82.9 tok/s
342 W43°CQ4_K_M
✓ Measured
Qwen2.5-Coder 32B82.9 tok/s
316 W43°CQ4_K_M
✓ Measured
Qwen2.5-Coder 32B (Q3_K_M)74.28 tok/s
482 W43°CQ3_K_M
✓ Measured
Qwen2.5-72B48.28 tok/s
553 W45°CQ4_K_M
✓ Measured
Meta-Llama-3.1-70B48.12 tok/s
550 W45°CQ4_K_M
✓ Measured
DeepSeek-R1-Distill-Llama-70B48.11 tok/s
551 W46°CQ4_K_M
✓ Measured
Hermes-4-70B48.1 tok/s
529 W46°CQ4_K_M
✓ Measured
Llama-3.3-70B-Instruct-abliterated48.1 tok/s
545 W46°CQ4_K_M
✓ Measured
Llama 3.3 70B47.97 tok/s
40.2 GB peak378 W46°C0.13 tok/WQ4_K_M
✓ Measured
Laguna-XS-2.1✕ Won't fit needs ~26 GBVRAM-gated at this precision✓ Measured
Nanbeige4.2-3B✕ Won't fit needs ~4 GBVRAM-gated at this precision✓ Measured

Image Generation images/min 5

FLUX.1 Schnell59.54
Z-Image Turbo (1024px)48.63
Z-Image Turbo39.675
Stable Diffusion XL29.2
FLUX.1 dev20.807
WorkloadResultTelemetryData
FLUX.1 Schnell59.54 images/min
619 W45°C
✓ Measured
Z-Image Turbo (1024px)48.63 images/min
776 W46°C
✓ Measured
Z-Image Turbo39.68 images/min
26.1 GB peak932 W50°C1.5 s/img
✓ Measured
Stable Diffusion XL29.2 images/min
16.5 GB peak488 W41°C2.1 s/img
✓ Measured
FLUX.1 dev20.81 images/min
37 GB peak989 W54°C2.9 s/img
✓ Measured

Fine-Tuning train tok/s 4

TinyLlama 1.1B LoRA26938.2
SmolLM2 1.7B LoRA26114.3
Qwen2.5 1.5B LoRA22351.1
Qwen2.5 7B LoRA14323.1
WorkloadResultTelemetryData
TinyLlama 1.1B LoRA26938.2 train tok/s
323 W37°C
✓ Measured
SmolLM2 1.7B LoRA26114.3 train tok/s
404 W40°C
✓ Measured
Qwen2.5 1.5B LoRA22351.1 train tok/s
371 W37°C
✓ Measured
Qwen2.5 7B LoRA14323.1 train tok/s
723 W45°C
✓ Measured

Image Editing images/min 2

WorkloadResultTelemetryData
FLUX.1 Kontext dev10.33 images/min
35.8 GB peak1030 W56°C5.8 s/img
✓ Measured
Qwen-Image-Edit8.14 images/min
60.5 GB peak1020 W55°C7.4 s/img
✓ Measured

Video Generation frames/s 2

WorkloadResultTelemetryData
LTX-Video (distilled)31.87 frames/s
60.5 GB peak716 W51°C3.2 s/clip
✓ Measured
Wan 2.2 5B (720p)2.94 frames/s
37 GB peak1013 W57°C16.7 s/clip
✓ Measured

Depth Estimation images/min 2

WorkloadResultTelemetryData
Depth Anything V2 Small1491.6 images/min
235 W35°C
✓ Measured
Depth Anything V2 Large1069.95 images/min
235 W35°C
✓ Measured

Segmentation images/min 2

WorkloadResultTelemetryData
SAM ViT-Base1719.59 images/min
235 W35°C
✓ Measured
SAM ViT-Huge399.37 images/min
278 W40°C
✓ Measured

Background Removal images/min 1

WorkloadResultTelemetryData
BiRefNet1064.73 images/min
240 W35°C
✓ Measured

Upscaling images/min 1

WorkloadResultTelemetryData
Swin2SR 4x Upscaler58.29 images/min
437 W42°C
✓ Measured

Extended workloads

Everything measured on this card beyond the standard 12-workload suite, grouped by the kind of work. Each group carries one unit, so numbers inside a group compare and numbers across groups do not.

Text Generation (tok/s)

GPT-OSS 20B346.76
DeepSeek-R1 Distill 8B289.05
Phi-4 14B178.74
Qwen3 14B168.03
Mistral Small 24B120.55
Gemma 3 27B93.31
WorkloadResultTelemetry
GPT-OSS 20B346.76 tok/s
11.9 GB peak286 W40°C1.22 tok/WQ4_K_M
DeepSeek-R1 Distill 8B289.05 tok/s
5.3 GB peak348 W42°C0.83 tok/WQ4_K_M
Phi-4 14B178.74 tok/s
9.2 GB peak392 W44°C0.46 tok/WQ4_K_M
Qwen3 14B168.03 tok/s
9 GB peak366 W43°C0.46 tok/WQ4_K_M
Mistral Small 24B120.55 tok/s
14 GB peak352 W44°C0.34 tok/WQ4_K_M
Gemma 3 27B93.31 tok/s
16.9 GB peak353 W43°C0.27 tok/WQ4_K_M

Image Generation (it/s)

Stable Diffusion 1.533.11
Qwen-Image9.08
FLUX.1 schnell7.75
WorkloadResultTelemetry
Stable Diffusion 1.533.11 it/s
26.7 GB peak361 W38°C0.9 s/img
Qwen-Image9.08 it/s
60.3 GB peak945 W50°C3.3 s/img
FLUX.1 schnell7.75 it/s
37 GB peak890 W48°C0.5 s/img
How we measured this. Every result comes from our own pinned, reproducible AI suite, 12 workloads: the Qwen3-4B to Llama-70B LLM ladder (llama.cpp, Q4_K_M), SDXL / Z-Image / FLUX-dev generation, FLUX-Kontext / Qwen-Edit editing, and LTX / Wan video, run first-party on rented hardware with under 0.5% run-to-run variance. Peak VRAM, power draw, temperature and tokens-per-watt are captured per workload. “Won’t fit” rows are real data: where a model exceeds the card’s VRAM at the tested precision we record a hard gate rather than silently dropping to a smaller quant. Measured 2026-07-12 · harness 2.1.0-b300-cuda13.

NVIDIA B300 specifications

ArchitectureBlackwell Ultra
CUDA cores20,480
VRAM288GB HBM3e
Memory bus8192-bit
Memory bandwidth8 TB/s
TDP1400 W
ProcessTSMC 4NP
InterfaceSXM6
Release date2025-11-01
Launch MSRP$40,000

Verdict, nothing in our suite slows it down

NVIDIA B300 scores 93.8/100, #1 of 102. It ran all 12 workloads. Every figure here is our own measurement.

Relative performance: where the NVIDIA B300 lands

100% = this card, AI & Machine Learning headline metric (AI Score). #1 of 21 desktop cards in this vertical.

GPURelative%AI Score
NVIDIA B300
100%93.8
NVIDIA B200
83%78
NVIDIA B100
71%67
NVIDIA H100 NVL
71%67
NVIDIA GH200 Grace Hopper
70%65.6

← All AI & Machine Learning GPU rankings

What this card can build

Whole-job timings, composed from our measured per-model results on this card.

WorkflowTimeEnergyBasis
50-image depth pass5 s0.18 Whmeasured
500-image masking run78 s5.8 Whmeasured
24-frame storyboard1.5 min11 Whall 2 stages measured
6-panel comic page2 min15.52 Whall 3 stages measured
60-second AI short film2.2 min14.29 Whall 3 stages measured
Character sheet, 12 poses2.2 min20.73 Whall 2 stages measured
200-product catalogue cutout3.7 min25.74 Whall 2 stages measured
10 short social clips4.7 min51.37 Whall 3 stages measured
40-product photo shoot6.1 min77.6 Whall 2 stages measured
Full codebase review6.3 min38.61 Whmeasured
40-product shoot, start to finish6.9 min82.75 Whall 4 stages measured
20 long-form articles9.7 min61.26 Whmeasured
100-photo restoration batch10.2 min166.12 Whmeasured
100-photo restore and enlarge11.9 min178.61 Whall 2 stages measured

Rent or buy?

This card is $40,000 to buy. The cheapest listed rate on RunPod is $6.940/hour, but that is the floor: we budget $8.328/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 4,803 GPU-hours. Below it you are paying for idle silicon.

How you would use itGPU-hours a yearRental cost a yearTime to break even
2 hours a day, hobby730$6,0796.6 years
8 hours a day, working on it2,920$24,3181.6 years
24/7, always-on agent8,760$72,9536.6 months

At hobby usage this card is very unlikely to pay for itself before it is superseded. Rent it. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.

Rental price

$6.940/hr+0.0% since 2026-08-14low $6.940 · high $6.940

Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.