Rental & access · measured on rented hardware · Updated July 2026
Where Can You Rent an NVIDIA B300 GPU?
B300s are rented by the hour from GPU cloud providers, and for practically everyone that's the only sensible way to touch one. The alternative is roughly $53,000 for a 1,400W module you have nowhere to plug in. We rented one and put it through 21 workloads for a few dollars. This page is what we found: what a rented B300 delivers, and, just as important, when renting one is a waste of your money.
288GB of HBM3e at 8 TB/s. We ran it through 21 workloads on 2026-07-12: 47.97 tok/s on Llama 3.3 70B, 9.71 it/s on FLUX.1-dev, 720p Wan video at 2.94 frames/s. The only card on our board of 102 where nothing in the core suite fails to fit.
Pros
47.97 tok/s on Llama 3.3 70B (measured, single stream)
288GB HBM3e, 12/12 core workloads fit, zero offload
8 TB/s bandwidth, the highest on our board
AI Score 93.8/100, the highest score on our board
Cons
Roughly $53,000 per GPU; no retail channel exists
1,400W rating, but we never measured above 1057.7W
Single-stream LLM work runs it at 12.2-40.3% utilisation
Measured on a CUDA 13 stack, unlike the rest of our fleet
Best for: Understanding what one B300 actually does before you spec, rent or buy anything built from them.
141GB and a measured 42.66 tok/s on Llama 3.3 70B, 89% of the B300's 47.97 on the same single-stream test, and all 12 core workloads fit here too. If your model fits in 141GB, the B300's extra memory is money you aren't using.
Pros
42.66 tok/s on Llama 3.3 70B, 89% of a B300 (measured)
All 12 core workloads fit in 141GB
Substantially cheaper per hour than Blackwell Ultra
AI Score 65.0/100
Cons
141GB vs 288GB. The ceiling is real if your model is enormous
Hopper generation; no FP4 Transformer Engine
700W board rating
Best for: The large majority of 70B-class inference work, where the B300's extra 147GB buys you nothing.
The workhorse. We measured 41.0 tok/s on Llama 3.3 70B on 80GB: 85% of a B300's single-stream speed, on hardware that's been in every cloud catalogue for years and prices accordingly. All 12 workloads fit.
Pros
41.0 tok/s on Llama 3.3 70B (measured), 85% of a B300
80GB holds everything in our core suite
The most widely available rental on the market
AI Score 62.6/100
Cons
80GB is the tightest of the three for a 70B
Hopper generation
700W board rating
Best for: Anyone who wants 70B inference working today at the lowest hourly rate that still fits the model.
You cannot meaningfully buy a B300. It costs roughly $53,000, it's an SXM6 module rated at 1,400W, and there is no slot in your computer that it fits, no power supply you own that will feed it, and no cooler on the market that will move a kilowatt of heat out of a tower case. You can, however, rent one for the price of a sandwich. That's what we did, and this entire page exists because of it: 21 workloads on rented Blackwell Ultra, a few dollars of compute, and a dataset nobody else has.
Everything a rented B300 ran, all 11 language models, single stream, Q4_K_M
Qwen3 4B
333.34 tok/s
Llama 3.1 8B
287.23 tok/s
Qwen2.5-Coder 14B
158.49 tok/s
Qwen3 32B
83.68 tok/s
Llama 3.3 70B
47.97 tok/s
GPT-OSS 20B
346.76 tok/s
Qwen3 14B
168.03 tok/s
Phi-4 14B
178.74 tok/s
Gemma 3 27B
93.31 tok/s
Mistral Small 24B
120.55 tok/s
DeepSeek-R1 Distill 8B
289.05 tok/s
Plus 5 image models and 2 video models. Not one gate, not one offload, not one compromise across the whole 21-workload run.
Llama 3.3 70B, single stream, what you actually gain from Blackwell Ultra
B300 (288GB)
47.97 tok/s
H200 (141GB)
42.66 tok/s
H100 (80GB)
41 tok/s
RTX PRO 6000 WS (96GB)
34.87 tok/s
All four fit the model. The B300 is 12% faster than an H200 and 17% faster than an H100, on hardware that costs a fraction of the hourly rate. Token generation is bandwidth-bound; once a model fits, the differences compress hard.
Why so close? Because token generation is bound by memory bandwidth, not compute, and once a model fits, the differences compress. The B300's 8 TB/s is a genuine lead over the H200's 4.8 TB/s, but it doesn't translate into a proportional lead on a single stream, because a single stream can't use it. Our power telemetry shows exactly that, and it's the most useful thing in this dataset.
377.8W
B300 running Llama 3.3 70B
27% of its 1,400W rating
31.3%
GPU utilisation on that run
the rest of the chip is idle
1029.5W
B300 running FLUX.1 Kontext
99% utilisation, saturated
3.6×
Power swing, same chip
the workload decides, not the card
Our verdict
Rent, don't buy: a $53,000 1,400W SXM6 module isn't going in your machine, and hourly rental put this entire 21-workload dataset on the page for a few dollars. But be deliberate about which card. Our measurements put an H200 at 42.66 tok/s and an H100 at 41.0 tok/s on Llama 3.3 70B against the B300's 47.97, 85-89% of the speed at a fraction of the rate, with all 12 core workloads fitting on every one of them. The B300 earns its price on models past 141GB, on long context, and on heavy batching. For a single-stream 70B, it's selling you 12%.
FAQ
Can you rent an NVIDIA B300?
Yes. B300 capacity is available on-demand from GPU cloud providers, typically with per-hour or per-minute billing and no long-term commitment. You can rent a single B300 or a full 8-GPU node, and some providers will provision full GB300 NVL72 racks on request. This is how we benchmarked ours, the entire 21-workload campaign cost a few dollars.
How much does it cost to rent a B300 per hour?
Rates move constantly with supply, provider, region and commitment, so we deliberately don't hardcode a number, it'd be wrong within a quarter and we'd rather point you at the live rate. As a concrete data point: our full 21-workload run on a rented B300, including every LLM up to 70B, five image models and two video models with telemetry throughout, cost a few dollars in total.
Should I rent a B300 or an H200?
Rent the H200 unless you specifically need more than 141GB. On our measurements the gap is much smaller than the spec sheets suggest: single-stream Llama 3.3 70B runs at 42.66 tok/s on an H200 versus 47.97 tok/s on a B300. The B300 is about 12% faster. Both fit all 12 core workloads. The B300's advantage is 288GB versus 141GB, plus FP4 and far better batched throughput. If your model fits in 141GB and you're not batching heavily, the extra you pay for Blackwell Ultra buys you 12%.
Is a B300 worth renting for a 70B model?
It works beautifully, 47.97 tok/s single stream, and the model's ~40.2GB footprint leaves over 240GB spare, but it's rarely the economical choice. We measured an H100 at 41.0 tok/s and an H200 at 42.66 tok/s on the same model and test, which is 85-89% of the B300's speed on much cheaper hardware. A B300 earns its rate on models too large for 141GB, on very long context where KV cache is the constraint, or on heavily batched serving where its FP4 throughput and memory headroom actually get used.
What can a B300 run that other GPUs can't?
In our suite, nothing, and that's the point. The B300 is the only card on our board of 102 where every one of the 12 core workloads fits with no offload and no compromise. But an H100 at 80GB also fits all 12. The difference isn't what runs, it's what's left over: our heaviest workload, Qwen-Image-Edit, peaked at 60.5GB on the B300, leaving over 225GB unused. That headroom is for models, context lengths and batch sizes bigger than anything in a standard inference suite. If you can't name the specific thing you need 288GB for, you probably don't need 288GB.
Why rent instead of buy?
A B300 costs roughly $53,000, is an SXM6 module rated at 1,400W, and cannot physically go in a normal machine. Beyond that, our telemetry makes the ownership case harder rather than easier: across 21 workloads the B300 never drew more than 1057.7W of its rating, and language-model inference averaged 286.2-391.6W at 12.2-40.3% utilisation. Owned hardware only pays for itself when it's saturated, and single-stream work doesn't come close to saturating this chip.
Is the B300 worth it over an H200, or even a 5090?
Depends on the workload, and I say this having rented all three. For pure text generation a 5090 performs surprisingly close to the B300: the giant card's edge really shows on image and video, the VRAM- and tensor-heavy work. The H200 is the sweet middle ground: it holds everything in our suite including the 70B class at a far friendlier hourly rate. Rent the B300 when the job is heavy visual generation or you need the capacity; rent the H200 for big LLM work; and if it's text on models that fit 32GB, the 5090 in your own desktop is closer than you'd believe.
How we test
Every number on this page describing a single B300 is our own measurement. We rented a B300 and ran the GPU Battle AI Suite v2 across 21 workloads, the core 12 that every GPU on this site runs, plus 9 Blackwell Ultra extras. LLMs run on llama.cpp (llama-bench) at Q4_K_M with -p 512 -n 128; diffusion and video run on diffusers/ComfyUI at BF16 (SDXL at FP16). Each workload gets a warmup pass plus multiple timed runs (5 for small LLMs, 3 for large models and images, 2 for video). We publish the mean as the result and the minimum as the 1% low. Run-to-run variance across our fleet is under 0.5%. Telemetry is sampled at 1 Hz from nvidia-smi for the duration of every run: power, temperature, utilisation, clocks and peak VRAM. Every wattage, temperature and tokens-per-watt figure on this page is logged draw, not a board rating. The B300 was measured on 2026-07-12 on a CUDA 13 stack (harness 2.1.0-b300-cuda13, driver 580.95.05, torch 2.13.0+cu130), which differs from the CUDA 12.8 stack the rest of our fleet runs. We label it rather than hide it. The caveat that matters most: these are single-GPU, single-stream, batch-size-1 numbers. That is the honest way to measure what one chip does, and it is deliberately not how a datacenter runs a B300. Vendor and MLPerf figures use large batches across many GPUs and will be far higher. Neither is wrong, they answer different questions. Ours answers 'what does one of these actually do'. System-level specifications (GPU counts, memory totals, rack power) come from NVIDIA's published documentation and are cited as such; we have not taken a rack apart. Where NVIDIA's own published system figures disagree with each other, we show the arithmetic rather than pick a side.