Rentals · datacenter GPUs · 6 cards measured · Updated October 2026
Newer datacenter GPUs are faster, but rarely by as much as they cost. We ran the same language, image and video models on six cards you can rent by the hour and priced each job at the cheapest rates we track. The fastest card is almost never the cheapest one per job.
Same jobs on six datacenter cards
| GPU | VRAM | Cheapest rent | Qwen3 32B tok/s | Llama 3.3 70B tok/s | FLUX.1-schnell img/min | LTX-Video frames/s | Per 1M tokens (70B) | Per 1,000 images |
|---|---|---|---|---|---|---|---|---|
| A100 80GB | 80GB | $0.95/hr | 45.5 | 24.4 | 28.4 | 8.9 | $10.78 | $0.56 |
| RTX PRO 6000 | 96GB | $1.08/hr | 70.1 | 34.9 | 42.4 | 16.4 | $8.57 | $0.42 |
| H100 | 80GB | $2.14/hr | 74.1 | 41.0 | 58.4 | 17.5 | $14.47 | $0.61 |
| H200 | 141GB | $3.59/hr | 76.6 | 42.7 | 60.5 | 17.9 | $23.38 | $0.99 |
| B200 | 192GB | $5.98/hr | 78.6 | 44.5 | 81.4 | 26.8 | $37.29 | $1.22 |
| B300 | 288GB | $6.94/hr | 83.7 | 48.0 | 59.5 | 31.9 | $40.19 | $1.94 |
Our own measurements; LLMs are single-stream llama.cpp at Q4_K_M. Costs are the cheapest hourly rate we tracked on RunPod or Vast.ai in October 2026, divided by the measured speed.
Cost per 1M tokens, Llama 3.3 70B
Cost per 1,000 images, FLUX.1-schnell
For language models, newer barely helps one user. Single-stream generation is limited by memory bandwidth and by how fast llama.cpp can feed the card, so Qwen3 32B goes from 74.1 tok/s on an H100 to 83.7 on a B300. That's 13% more speed for 3.2 times the price. Per million tokens on Llama 3.3 70B, the RTX PRO 6000 is cheapest at $8.57, and the B300 is the most expensive at $40.19.
The RTX PRO 6000 is the sleeper. It's a 96GB Blackwell workstation card that rents for about $1.08 an hour, and on gpt-oss-120b it ran at 254 tok/s, faster than the H100's 220. For models between 80 and 96GB it's the cheapest single card that fits.
For image and video, Blackwell pulls ahead. Diffusion is compute-bound, so the B200 makes 81.4 FLUX.1-schnell images a minute against the H100's 58.4. Oddly, the B300 makes only 59.5: its extra transistors went to low-precision AI math that this kind of model doesn't use, which we dig into in our B300-vs-B200 article. Per 1,000 images the RTX PRO 6000 is cheapest at $0.42.
Memory decides the giant models. Qwen3 235B needs about 160GB at Q4, so it runs on a B200 (80.2 tok/s) or a B300 (90.6 tok/s) and nothing smaller. GLM-5.3-Flash (72.7 tok/s) and DeepSeek-V3.2 at 2-bit (50.4 tok/s) need more than 200GB, and of the cards we track, only the B300's 288GB holds them on one GPU.
Rent by memory first and price per job second. For language models the RTX PRO 6000 is cheapest per token, while the B300 is only 13% faster than an H100 on Qwen3 32B at 3.2x the price. The B200 leads image and video, the RTX PRO 6000 is the cheapest 96GB card, and the B300 earns its price only for models over 192GB.
Every number is our own measurement on rented cards: LLMs with llama.cpp llama-bench at Q4_K_M (single stream), images with diffusers in bf16 (three timed generations after a warmup), video with diffusers (frames of output per second). Rental rates are the cheapest hourly prices we tracked on RunPod and Vast.ai in October 2026; the rentals page has live numbers.