288GB · AI Score 93.8/100 · first-party measured on 12 AI workloads
93.8 AI Score ✓ Measured
Every number on this page is first-party: NVIDIA B300 was run on our pinned 12-workload AI suite on 2026-07-12, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) NVIDIA B300 delivers about 287.23 tokens/sec. Stepping up to Qwen3 32B it holds roughly 83.68 tok/s. The full Llama 3.3 70B still runs, at about 47.97 tok/s. For image generation, SDXL runs at 14.6 it/s, and FLUX.1-dev at 9.71 it/s. All 12 workloads fit in 288GB. There is no model in our suite this card has to turn down. NVIDIA B300 isn't a retail purchase for most people. It's rented by the hour. You can run this exact card on RunPod.
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| LFM2.5-1.2B | 955.13 tok/s | 259 W38°CQ4_K_M | ✓ Measured |
| gemma-3-270m | 926.67 tok/s | 242 W33°CQ4_K_M | ✓ Measured |
| Qwen1.5-0.5B | 894.84 tok/s | 246 W34°CQ4_K_M | ✓ Measured |
| Llama 3.2 1B | 889.19 tok/s | 263 W44°CQ4_K_M | ✓ Measured |
| SmolLM2-135M | 855.87 tok/s | 244 W34°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder-0.5B | 843.44 tok/s | 244 W33°CQ4_K_M | ✓ Measured |
| Qwen2.5-0.5B | 843.4 tok/s | 252 W33°CQ4_K_M | ✓ Measured |
| Qwen3-0.6B | 742.64 tok/s | 249 W34°CQ4_K_M | ✓ Measured |
| Qwen3 0.6B | 718.58 tok/s | 223 W42°CQ4_K_M | ✓ Measured |
| LFM2.5-8B-A1B | 596.16 tok/s | 277 W36°CQ4_K_M | ✓ Measured |
| Qwen2.5-1.5B | 584.96 tok/s | 267 W35°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder-1.5B | 584.5 tok/s | 265 W37°CQ4_K_M | ✓ Measured |
| Qwen2-1.5B | 583.79 tok/s | 265 W36°CQ4_K_M | ✓ Measured |
| Qwen3-1.7B | 575.97 tok/s | 274 W37°CQ4_K_M | ✓ Measured |
| DeepSeek-R1 Distill 1.5B | 562.14 tok/s | 274 W44°CQ4_K_M | ✓ Measured |
| Qwen3 1.7B | 555.1 tok/s | 275 W44°CQ4_K_M | ✓ Measured |
| gemma-3-1b | 533.32 tok/s | 271 W35°CQ4_K_M | ✓ Measured |
| Llama-3.2-3B-Instruct-uncensored | 448.94 tok/s | 323 W39°CQ4_K_M | ✓ Measured |
| Hermes-3-Llama-3.2-3B | 446.78 tok/s | 257 W37°CQ4_K_M | ✓ Measured |
| Llama 3.2 3B | 432.24 tok/s | 268 W46°CQ4_K_M | ✓ Measured |
| SmolLM3-3B | 420.3 tok/s | 321 W37°CQ4_K_M | ✓ Measured |
| gemma-2-2b-it-abliterated | 418.18 tok/s | 305 W37°CQ4_K_M | ✓ Measured |
| gemma-2-2b | 417.84 tok/s | 307 W36°CQ4_K_M | ✓ Measured |
| SmolLM3 3B | 411.25 tok/s | 278 W46°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder-3B | 410.06 tok/s | 317 W38°CQ4_K_M | ✓ Measured |
| Phi-4-mini | 409.58 tok/s | 324 W39°CQ4_K_M | ✓ Measured |
| Qwen2.5-3B | 409.56 tok/s | 300 W37°CQ4_K_M | ✓ Measured |
| Phi-4 Mini 3.8B | 398.16 tok/s | 301 W47°CQ4_K_M | ✓ Measured |
| AI21-Jamba-Reasoning-3B | 380.88 tok/s | 311 W37°CQ4_K_M | ✓ Measured |
| phi-2 | 356.65 tok/s | 316 W37°CQ4_K_M | ✓ Measured |
| gpt-oss-20b | 355.56 tok/s | 279 W37°CQ4_K_M | ✓ Measured |
| Phi-3.5-mini | 352.25 tok/s | 334 W37°CQ4_K_M | ✓ Measured |
| Nemotron-3-Nano-30B-A3B | 345.31 tok/s | 280 W37°CQ4_K_M | ✓ Measured |
| DeepSeek-Coder-V2-Lite | 340.53 tok/s | 279 W37°CQ4_K_M | ✓ Measured |
| Qwen3-4B-Instruct-2507 | 339.78 tok/s | 325 W38°CQ4_K_M | ✓ Measured |
| Qwen3-4B-Thinking-2507 | 339.74 tok/s | 325 W37°CQ4_K_M | ✓ Measured |
| Qwen3-4B-Instruct-2507 | 339.72 tok/s | 313 W38°CQ4_K_M | ✓ Measured |
| Qwen3 4B | 333.34 tok/s | 3.1 GB peak295 W39°C1.13 tok/WQ4_K_M | ✓ Measured |
| Llama-2-7B | 315.9 tok/s | 379 W40°CQ4_K_M | ✓ Measured |
| Mistral-7B-Instruct-v0.2 | 306.79 tok/s | 377 W40°CQ4_K_M | ✓ Measured |
| Mistral-7B-Instruct-v0.3 | 306.46 tok/s | 366 W40°CQ4_K_M | ✓ Measured |
| Mistral-7B-Instruct-v0.1 | 306.44 tok/s | 372 W40°CQ4_K_M | ✓ Measured |
| Gemma 3 4B | 301.07 tok/s | 283 W46°CQ4_K_M | ✓ Measured |
| Mistral 7B v0.3 | 298.95 tok/s | 313 W48°CQ4_K_M | ✓ Measured |
| Qwen2.5-7B | 293.63 tok/s | 370 W39°CQ4_K_M | ✓ Measured |
| L3-8B-Stheno-v3.2 | 293 tok/s | 381 W40°CQ4_K_M | ✓ Measured |
| Meta-Llama-3.1-8B | 293 tok/s | 352 W40°CQ4_K_M | ✓ Measured |
| DarkIdol-Llama-3.1-8B-Instruct-1.2-Uncensored | 292.55 tok/s | 361 W41°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder-7B-Instruct-abliterated | 292.21 tok/s | 370 W39°CQ4_K_M | ✓ Measured |
| Qwen3-30B-A3B | 291.82 tok/s | 287 W36°CQ4_K_M | ✓ Measured |
| dolphin-2.9-llama3-8b | 290.47 tok/s | 354 W40°CQ4_K_M | ✓ Measured |
| Llama 3.1 8B | 287.23 tok/s | 5.3 GB peak338 W41°C0.85 tok/WQ4_K_M | ✓ Measured |
| Dolphin X1 8B | 287.02 tok/s | 292 W40°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder 7B | 286.5 tok/s | 299 W48°CQ4_K_M | ✓ Measured |
| Dolphin 3.0 Llama 3.1 8B | 285.88 tok/s | 300 W40°CQ4_K_M | ✓ Measured |
| Qwen3-Coder 30B A3B | 284.88 tok/s | 256 W36°CQ4_K_M | ✓ Measured |
| DeepSeek-R1 Distill 7B | 284.77 tok/s | 291 W42°CQ4_K_M | ✓ Measured |
| Qwen3 30B A3B | 284.66 tok/s | 250 W39°CQ4_K_M | ✓ Measured |
| DeepSeek-R1 Distill Llama 8B | 283.77 tok/s | 306 W42°CQ4_K_M | ✓ Measured |
| DeepSeek-R1-0528-Qwen3-8B | 270.35 tok/s | 375 W40°CQ4_K_M | ✓ Measured |
| Qwen3-8B | 270.35 tok/s | 383 W41°CQ4_K_M | ✓ Measured |
| Josiefied-Qwen3-8B-abliterated-v1 | 270.09 tok/s | 362 W40°CQ4_K_M | ✓ Measured |
| Qwen3 30B A3B (Q3_K_M) | 268.71 tok/s | 302 W37°CQ3_K_M | ✓ Measured |
| Qwen3 8B | 261.4 tok/s | 275 W41°CQ4_K_M | ✓ Measured |
| Dolphin X1 Trinity Nano 6B | 257.93 tok/s | 249 W33°CQ4_K_M | ✓ Measured |
| Ornith-1.0-9B | 240.67 tok/s | 377 W40°CQ4_K_M | ✓ Measured |
| KAT-Coder-V2.5-Dev | 239.66 tok/s | 287 W36°CQ4_K_M | ✓ Measured |
| Qwen-AgentWorld-35B-A3B | 233.04 tok/s | 293 W36°CQ4_K_M | ✓ Measured |
| Ornith-1.0-35B | 232.36 tok/s | 291 W36°CQ4_K_M | ✓ Measured |
| gemma-2-9b | 202.22 tok/s | 386 W40°CQ4_K_M | ✓ Measured |
| Qwen3-Next-80B-A3B-Thinking | 200.55 tok/s | 287 W36°CQ4_K_M | ✓ Measured |
| Qwen3-Coder-Next | 199.39 tok/s | 288 W36°CQ4_K_M | ✓ Measured |
| Qwen3-Coder-Next-abliterated | 199.12 tok/s | 280 W36°CQ4_K_M | ✓ Measured |
| NemoMix-Unleashed-12B | 197.06 tok/s | 403 W41°CQ4_K_M | ✓ Measured |
| Mistral-Nemo-Instruct-2407 | 196.92 tok/s | 406 W41°CQ4_K_M | ✓ Measured |
| GLM-4.7-Flash | 196.52 tok/s | 303 W35°CQ4_K_M | ✓ Measured |
| Qwen3-Coder-Next | 196.34 tok/s | 292 W36°CQ4_K_M | ✓ Measured |
| Qwen3-Next-80B-A3B-Thinking | 194.98 tok/s | 292 W36°CQ4_K_M | ✓ Measured |
| Qwen3-Next-80B-A3B | 191.32 tok/s | 284 W36°CQ4_K_M | ✓ Measured |
| GLM-4.7-Flash-REAP-23B-A3B | 178.49 tok/s | 300 W36°CQ4_K_M | ✓ Measured |
| Phi-4 14B | 177.41 tok/s | 334 W44°CQ4_K_M | ✓ Measured |
| Qwen3-14B | 169.52 tok/s | 422 W41°CQ4_K_M | ✓ Measured |
| Hermes-4-14B | 169.51 tok/s | 422 W41°CQ4_K_M | ✓ Measured |
| Qwen3 14B | 165.44 tok/s | 315 W43°CQ4_K_M | ✓ Measured |
| Phi-4 14B (Q3_K_M) | 165.06 tok/s | 439 W42°CQ3_K_M | ✓ Measured |
| Gemma 3 12B | 161.27 tok/s | 315 W42°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder-14B-Instruct-abliterated | 159.5 tok/s | 423 W40°CQ4_K_M | ✓ Measured |
| Uncensored | 159.48 tok/s | 416 W41°CQ4_K_M | ✓ Measured |
| EVA-Qwen2.5-14B-v0.2 | 159.46 tok/s | 416 W41°CQ4_K_M | ✓ Measured |
| Qwen2.5-14B | 159.33 tok/s | 424 W40°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder 14B | 158.49 tok/s | 9 GB peak367 W42°C0.43 tok/WQ4_K_M | ✓ Measured |
| DeepSeek-R1 Distill 14B | 156.33 tok/s | 328 W43°CQ4_K_M | ✓ Measured |
| Gemma 4 12B | 155.96 tok/s | 314 W42°CQ4_K_M | ✓ Measured |
| Gemma 3 12B (Q3_K_M) | 149.61 tok/s | 412 W40°CQ3_K_M | ✓ Measured |
| DeepSeek-R1 Distill 14B (Q3_K_M) | 144.74 tok/s | 424 W41°CQ3_K_M | ✓ Measured |
| StarCoder2 15B | 144.05 tok/s | 330 W41°CQ4_K_M | ✓ Measured |
| Cydonia-24B-v4.3 | 121.9 tok/s | 477 W43°CQ4_K_M | ✓ Measured |
| Dolphin-Mistral-24B-Venice-Edition | 121.87 tok/s | 473 W43°CQ4_K_M | ✓ Measured |
| Devstral Small 24B | 121.18 tok/s | 323 W42°CQ4_K_M | ✓ Measured |
| Dolphin 3.0 R1 Mistral 24B | 120.95 tok/s | 348 W43°CQ4_K_M | ✓ Measured |
| Dolphin Mistral 24B Venice | 120.93 tok/s | 331 W43°CQ4_K_M | ✓ Measured |
| Codestral 22B | 119.81 tok/s | 328 W42°CQ4_K_M | ✓ Measured |
| Mistral Small 24B | 119.3 tok/s | 345 W45°CQ4_K_M | ✓ Measured |
| Mistral Small 24B (Q3_K_M) | 108.37 tok/s | 460 W43°CQ3_K_M | ✓ Measured |
| Codestral 22B (Q3_K_M) | 107.98 tok/s | 465 W43°CQ3_K_M | ✓ Measured |
| Gemma 3 27B | 93.44 tok/s | 326 W42°CQ4_K_M | ✓ Measured |
| Olmo-3.1-32B-Think | 87.85 tok/s | 476 W43°CQ4_K_M | ✓ Measured |
| Dolphin 2.9.1 Yi 1.5 34B | 85.15 tok/s | 312 W43°CQ4_K_M | ✓ Measured |
| Qwen3 32B | 83.68 tok/s | 19.1 GB peak346 W44°C0.24 tok/WQ4_K_M | ✓ Measured |
| Qwen2.5-32B | 83.41 tok/s | 490 W43°CQ4_K_M | ✓ Measured |
| DeepSeek-R1-Distill-Qwen-32B-abliterated | 83.19 tok/s | 468 W43°CQ4_K_M | ✓ Measured |
| QwQ 32B | 82.91 tok/s | 328 W43°CQ4_K_M | ✓ Measured |
| DeepSeek-R1 Distill 32B | 82.9 tok/s | 342 W43°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder 32B | 82.9 tok/s | 316 W43°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder 32B (Q3_K_M) | 74.28 tok/s | 482 W43°CQ3_K_M | ✓ Measured |
| Qwen2.5-72B | 48.28 tok/s | 553 W45°CQ4_K_M | ✓ Measured |
| Meta-Llama-3.1-70B | 48.12 tok/s | 550 W45°CQ4_K_M | ✓ Measured |
| DeepSeek-R1-Distill-Llama-70B | 48.11 tok/s | 551 W46°CQ4_K_M | ✓ Measured |
| Hermes-4-70B | 48.1 tok/s | 529 W46°CQ4_K_M | ✓ Measured |
| Llama-3.3-70B-Instruct-abliterated | 48.1 tok/s | 545 W46°CQ4_K_M | ✓ Measured |
| Llama 3.3 70B | 47.97 tok/s | 40.2 GB peak378 W46°C0.13 tok/WQ4_K_M | ✓ Measured |
| Laguna-XS-2.1 | ✕ Won't fit needs ~26 GB | VRAM-gated at this precision | ✓ Measured |
| Nanbeige4.2-3B | ✕ Won't fit needs ~4 GB | VRAM-gated at this precision | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| FLUX.1 Schnell | 59.54 images/min | 619 W45°C | ✓ Measured |
| Z-Image Turbo (1024px) | 48.63 images/min | 776 W46°C | ✓ Measured |
| Z-Image Turbo | 39.68 images/min | 26.1 GB peak932 W50°C1.5 s/img | ✓ Measured |
| Stable Diffusion XL | 29.2 images/min | 16.5 GB peak488 W41°C2.1 s/img | ✓ Measured |
| FLUX.1 dev | 20.81 images/min | 37 GB peak989 W54°C2.9 s/img | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| TinyLlama 1.1B LoRA | 26938.2 train tok/s | 323 W37°C | ✓ Measured |
| SmolLM2 1.7B LoRA | 26114.3 train tok/s | 404 W40°C | ✓ Measured |
| Qwen2.5 1.5B LoRA | 22351.1 train tok/s | 371 W37°C | ✓ Measured |
| Qwen2.5 7B LoRA | 14323.1 train tok/s | 723 W45°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| FLUX.1 Kontext dev | 10.33 images/min | 35.8 GB peak1030 W56°C5.8 s/img | ✓ Measured |
| Qwen-Image-Edit | 8.14 images/min | 60.5 GB peak1020 W55°C7.4 s/img | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| LTX-Video (distilled) | 31.87 frames/s | 60.5 GB peak716 W51°C3.2 s/clip | ✓ Measured |
| Wan 2.2 5B (720p) | 2.94 frames/s | 37 GB peak1013 W57°C16.7 s/clip | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Depth Anything V2 Small | 1491.6 images/min | 235 W35°C | ✓ Measured |
| Depth Anything V2 Large | 1069.95 images/min | 235 W35°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| SAM ViT-Base | 1719.59 images/min | 235 W35°C | ✓ Measured |
| SAM ViT-Huge | 399.37 images/min | 278 W40°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| BiRefNet | 1064.73 images/min | 240 W35°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Swin2SR 4x Upscaler | 58.29 images/min | 437 W42°C | ✓ Measured |
Everything measured on this card beyond the standard 12-workload suite, grouped by the kind of work. Each group carries one unit, so numbers inside a group compare and numbers across groups do not.
| Workload | Result | Telemetry |
|---|---|---|
| GPT-OSS 20B | 346.76 tok/s | 11.9 GB peak286 W40°C1.22 tok/WQ4_K_M |
| DeepSeek-R1 Distill 8B | 289.05 tok/s | 5.3 GB peak348 W42°C0.83 tok/WQ4_K_M |
| Phi-4 14B | 178.74 tok/s | 9.2 GB peak392 W44°C0.46 tok/WQ4_K_M |
| Qwen3 14B | 168.03 tok/s | 9 GB peak366 W43°C0.46 tok/WQ4_K_M |
| Mistral Small 24B | 120.55 tok/s | 14 GB peak352 W44°C0.34 tok/WQ4_K_M |
| Gemma 3 27B | 93.31 tok/s | 16.9 GB peak353 W43°C0.27 tok/WQ4_K_M |
| Workload | Result | Telemetry |
|---|---|---|
| Stable Diffusion 1.5 | 33.11 it/s | 26.7 GB peak361 W38°C0.9 s/img |
| Qwen-Image | 9.08 it/s | 60.3 GB peak945 W50°C3.3 s/img |
| FLUX.1 schnell | 7.75 it/s | 37 GB peak890 W48°C0.5 s/img |
| Architecture | Blackwell Ultra |
| CUDA cores | 20,480 |
| VRAM | 288GB HBM3e |
| Memory bus | 8192-bit |
| Memory bandwidth | 8 TB/s |
| TDP | 1400 W |
| Process | TSMC 4NP |
| Interface | SXM6 |
| Release date | 2025-11-01 |
| Launch MSRP | $40,000 |
NVIDIA B300 scores 93.8/100, #1 of 102. It ran all 12 workloads. Every figure here is our own measurement.
100% = this card, AI & Machine Learning headline metric (AI Score). #1 of 21 desktop cards in this vertical.
| GPU | Relative | % | AI Score |
|---|---|---|---|
| NVIDIA B300 | 100% | 93.8 | |
| NVIDIA B200 | 83% | 78 | |
| NVIDIA B100 | 71% | 67 | |
| NVIDIA H100 NVL | 71% | 67 | |
| NVIDIA GH200 Grace Hopper | 70% | 65.6 |
← All AI & Machine Learning GPU rankings
Whole-job timings, composed from our measured per-model results on this card.
| Workflow | Time | Energy | Basis |
|---|---|---|---|
| 50-image depth pass | 5 s | 0.18 Wh | measured |
| 500-image masking run | 78 s | 5.8 Wh | measured |
| 24-frame storyboard | 1.5 min | 11 Wh | all 2 stages measured |
| 6-panel comic page | 2 min | 15.52 Wh | all 3 stages measured |
| 60-second AI short film | 2.2 min | 14.29 Wh | all 3 stages measured |
| Character sheet, 12 poses | 2.2 min | 20.73 Wh | all 2 stages measured |
| 200-product catalogue cutout | 3.7 min | 25.74 Wh | all 2 stages measured |
| 10 short social clips | 4.7 min | 51.37 Wh | all 3 stages measured |
| 40-product photo shoot | 6.1 min | 77.6 Wh | all 2 stages measured |
| Full codebase review | 6.3 min | 38.61 Wh | measured |
| 40-product shoot, start to finish | 6.9 min | 82.75 Wh | all 4 stages measured |
| 20 long-form articles | 9.7 min | 61.26 Wh | measured |
| 100-photo restoration batch | 10.2 min | 166.12 Wh | measured |
| 100-photo restore and enlarge | 11.9 min | 178.61 Wh | all 2 stages measured |
This card is $40,000 to buy. The cheapest listed rate on RunPod is $6.940/hour, but that is the floor: we budget $8.328/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 4,803 GPU-hours. Below it you are paying for idle silicon.
| How you would use it | GPU-hours a year | Rental cost a year | Time to break even |
|---|---|---|---|
| 2 hours a day, hobby | 730 | $6,079 | 6.6 years |
| 8 hours a day, working on it | 2,920 | $24,318 | 1.6 years |
| 24/7, always-on agent | 8,760 | $72,953 | 6.6 months |
At hobby usage this card is very unlikely to pay for itself before it is superseded. Rent it. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.
Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.