Rentals · what we learned · from our own benchmark runs · Updated October 2026

Mistakes People Make Renting Cloud GPUs

Every number on GPU Battle comes from a GPU we rented by the hour, on Vast.ai, RunPod and Modal. That's more than 2,500 benchmark runs across hundreds of rentals, and plenty of them went wrong. These are the mistakes that cost us money or gave us bad numbers, what they look like, and how to avoid them.

10x
How slow one bad host ran
an RTX A5000 at 11 tok/s instead of about 120
230-315W
Power caps we found on rented RTX 3090s
stock is 350W or more
30min
Compiling llama.cpp on a 2016 Xeon
billed the whole time
5.5x
H100 vs RTX 5090, cost per token
Qwen3 8B, cheapest rates

1. Shopping by the hourly price. The hourly rate isn't the bill; the job is. Qwen3 8B runs at about the same speed on an RTX 5090 (244 tok/s) as on an H100 (244), but the 5090 rents for $0.39 an hour against $2.14, so a million tokens costs $0.44 instead of $2.43. Divide the rate by the speed for your model before you compare.

2. Not checking the power limit. On consumer hosts, power caps are the norm, not the exception. Rented RTX 3090s we checked were capped at 230 to 315W of their 350W-plus stock limit, 4090s at 320 to 420W of 450W, and one 5090 at 450 of 575W. A capped card is a slower card at the same price. Run nvidia-smi and look at the power limit in the first minute, and leave if it's well under stock.

3. Ignoring the CPU, RAM and disk. A cheap RTX 3060 we rented sat on a 2016 Xeon that took 30 minutes to compile llama.cpp, all of it billed. We now refuse hosts with fewer than 16 CPU cores, 32GB of RAM, 800Mb/s of internet or 500MB/s of disk. The GPU is only as fast as the machine feeding it.

4. Paying for downloads. A 40 to 60GB model takes a while to download, and on most platforms you pay the GPU rate the whole time. On Vast.ai the host also sets its own bandwidth price. For our image benchmarks on popular cards, the download was most of the roughly 40-cent cost per job. Download once to cheap storage (Modal lets you do it on a CPU machine first), or pick hosts with fast, cheap bandwidth.

5. Trusting the first number. One RTX A5000 host passed every check and still ran Qwen3 8B at 11 tok/s instead of about 120. A hot container read 12 to 29% low. Before you trust a host, run one small known model and compare against a reference like our tables; bad hosts show up in the first minute.

6. Picking a card with a tiny pool. Price doesn't predict whether a rental works; the number of healthy hosts does. When we last checked, Vast.ai had about 30 healthy RTX 4090 hosts and 29 RTX 5090s, but about 10 RTX 3090s and only 3 RTX 3060s, all on old Xeons. With a deep pool you can drop a bad host in a minute and try the next for about a cent.

7. Renting a card the model doesn't fit. A model that doesn't fit in VRAM doesn't run slowly, it doesn't run, and you pay for the setup either way. Check the model's measured VRAM first; every model page here lists it, and the giant models are strict: Qwen3 235B needs about 160GB, and GLM-5.3-Flash needs the B300's 288GB.

8. Forgetting load time on giant models. On a B300, loading a 150 to 250GB model took us 2.5 to 4.5 minutes before the first token, at about $7 an hour. Batch your work so you load once, and don't restart the container between jobs.

9. Leaving it running. The most expensive GPU is the one you forgot. We destroy every instance the moment a job finishes; if you're renting by the hour, set a timer or use a platform that shuts down when the job ends.

Before you rent: a one-minute checklist

CheckWhat to look forWhy
Cost per jobrate ÷ measured speed for your modelthe cheapest hour is often not the cheapest job
Power limitnvidia-smi: limit near the card's stock TDPcapped cards run slower at the same price
Host machine16+ cores, 32GB+ RAM, fast disk and internetslow hosts bill you for waiting
Bandwidth priceper-GB charges on marketplace hostsbig model downloads add up
VRAMthe model's measured peak vs the cardtoo small means it won't load at all
First resultone small known model vs a reference numberbad hosts show up in a minute
Shutdowndestroy, don't just stopstopped instances can still bill storage

Our verdict

Price the job, not the hour; check the power limit, the host machine and the bandwidth price before you start; test one small model against a known number; and destroy the instance when you're done. On Qwen3 8B a rented RTX 5090 costs $0.44 per million tokens against $2.43 on an H100, so the right card for your model matters more than the platform.

FAQ

Why is my rented GPU slower than benchmarks?
Usually a power cap, a weak host CPU or disk, thermal throttling, or a bad host. Check nvidia-smi for the power limit and run one small known model against a reference number before trusting the machine.
Is Vast.ai or RunPod cheaper?
Vast.ai usually has the lower hourly rate, but hosts set their own bandwidth and storage prices. For jobs that download big models or move a lot of data, RunPod's fixed pricing can work out cheaper.
How do I avoid paying while a model downloads?
Download to persistent storage once and reuse it, use a platform that lets you download on a CPU machine (Modal does), or choose hosts with fast internet and low bandwidth charges.
What should I check before renting a GPU?
Cost per job for your model, the card's power limit, the host's CPU, RAM, disk and internet, bandwidth pricing, whether your model fits in VRAM, and a quick test against a known number.

How we test

Everything here comes from the rentals behind GPU Battle's benchmarks on Vast.ai, RunPod and Modal in August to October 2026: host checks we log before every run (power limit, CPU cores, RAM, bandwidth, disk), runs we threw out as bad hosts, and the cheapest hourly rates we track. Cost-per-token figures use our measured Qwen3 8B speeds and October 2026 rates.