Rentals · datacenter GPUs · 6 cards measured · Updated October 2026

H100 vs H200 vs B200 vs B300

Newer datacenter GPUs are faster, but rarely by as much as they cost. We ran the same language, image and video models on six cards you can rent by the hour and priced each job at the cheapest rates we track. The fastest card is almost never the cheapest one per job.

+13%
B300 vs H100 on Qwen3 32B
83.7 vs 74.1 tok/s
3.2x
B300 vs H100 hourly price
$6.94 vs $2.14
RTX PRO 6000
Cheapest per token on Llama 3.3 70B
$8.57 per 1M tokens
288GB
B300 memory
the only single card for 200GB+ models

Same jobs on six datacenter cards

A100 80GB80GB
RTX PRO 600096GB
H10080GB
H200141GB
B200192GB
B300288GB
GPUVRAMCheapest rentQwen3 32B tok/sLlama 3.3 70B tok/sFLUX.1-schnell img/minLTX-Video frames/sPer 1M tokens (70B)Per 1,000 images
A100 80GB80GB$0.95/hr45.524.428.48.9$10.78$0.56
RTX PRO 600096GB$1.08/hr70.134.942.416.4$8.57$0.42
H10080GB$2.14/hr74.141.058.417.5$14.47$0.61
H200141GB$3.59/hr76.642.760.517.9$23.38$0.99
B200192GB$5.98/hr78.644.581.426.8$37.29$1.22
B300288GB$6.94/hr83.748.059.531.9$40.19$1.94

Our own measurements; LLMs are single-stream llama.cpp at Q4_K_M. Costs are the cheapest hourly rate we tracked on RunPod or Vast.ai in October 2026, divided by the measured speed.

Cost per 1M tokens, Llama 3.3 70B

A100 80GB
10.78 $ per 1M tokens
RTX PRO 6000
8.57 $ per 1M tokens
H100
14.47 $ per 1M tokens
H200
23.38 $ per 1M tokens
B200
37.29 $ per 1M tokens
B300
40.19 $ per 1M tokens

Cost per 1,000 images, FLUX.1-schnell

A100 80GB
0.56 $ per 1,000 images
RTX PRO 6000
0.42 $ per 1,000 images
H100
0.61 $ per 1,000 images
H200
0.99 $ per 1,000 images
B200
1.22 $ per 1,000 images
B300
1.94 $ per 1,000 images

For language models, newer barely helps one user. Single-stream generation is limited by memory bandwidth and by how fast llama.cpp can feed the card, so Qwen3 32B goes from 74.1 tok/s on an H100 to 83.7 on a B300. That's 13% more speed for 3.2 times the price. Per million tokens on Llama 3.3 70B, the RTX PRO 6000 is cheapest at $8.57, and the B300 is the most expensive at $40.19.

The RTX PRO 6000 is the sleeper. It's a 96GB Blackwell workstation card that rents for about $1.08 an hour, and on gpt-oss-120b it ran at 254 tok/s, faster than the H100's 220. For models between 80 and 96GB it's the cheapest single card that fits.

For image and video, Blackwell pulls ahead. Diffusion is compute-bound, so the B200 makes 81.4 FLUX.1-schnell images a minute against the H100's 58.4. Oddly, the B300 makes only 59.5: its extra transistors went to low-precision AI math that this kind of model doesn't use, which we dig into in our B300-vs-B200 article. Per 1,000 images the RTX PRO 6000 is cheapest at $0.42.

Memory decides the giant models. Qwen3 235B needs about 160GB at Q4, so it runs on a B200 (80.2 tok/s) or a B300 (90.6 tok/s) and nothing smaller. GLM-5.3-Flash (72.7 tok/s) and DeepSeek-V3.2 at 2-bit (50.4 tok/s) need more than 200GB, and of the cards we track, only the B300's 288GB holds them on one GPU.

Our verdict

Rent by memory first and price per job second. For language models the RTX PRO 6000 is cheapest per token, while the B300 is only 13% faster than an H100 on Qwen3 32B at 3.2x the price. The B200 leads image and video, the RTX PRO 6000 is the cheapest 96GB card, and the B300 earns its price only for models over 192GB.

FAQ

Is the B300 faster than the H100?
Yes, but not by much for one LLM user: 83.7 vs 74.1 tok/s on Qwen3 32B. It costs 3.2x as much per hour, so it's only worth it for models that need its 288GB.
Which datacenter GPU is cheapest for LLMs?
Per token on Llama 3.3 70B, the RTX PRO 6000 at $8.57 per million tokens, at October 2026 rental rates.
Can a single GPU run Qwen3 235B?
Yes, on a B200 (80.2 tok/s) or a B300 (90.6 tok/s) at Q4_K_M, which needs about 160GB. An H200's 141GB is not enough.
Why is the B300 slower than the B200 for image generation?
On FLUX.1-schnell the B300 made 59.5 images a minute against the B200's 81.4. The B300's extra silicon is aimed at low-precision math for language models, which bf16 diffusion doesn't use.

How we test

Every number is our own measurement on rented cards: LLMs with llama.cpp llama-bench at Q4_K_M (single stream), images with diffusers in bf16 (three timed generations after a warmup), video with diffusers (frames of output per second). Rental rates are the cheapest hourly prices we tracked on RunPod and Vast.ai in October 2026; the rentals page has live numbers.