Buyer's guide · local LLMs · 8 consumer cards measured · Updated October 2026
We ran the same language models on eight consumer cards, from the 12GB RTX 3060 to the 32GB RTX 5090. Two numbers decide which card is right for you: how much VRAM it has, which decides which models load at all, and how fast it generates once they do. Here's both, measured.
Generation speed on the same models (tok/s, Q4_K_M, llama.cpp)
| GPU | VRAM | Launch price | Qwen3 8B | Qwen3 14B | Qwen3 30B-A3B | Qwen3 32B | Cheapest rent |
|---|---|---|---|---|---|---|---|
| RTX 3060 | 12GB | $329 | 63.8 | 36.5 | doesn't fit | doesn't fit | $0.04/hr |
| RTX 4060 Ti 16GB | 16GB | $499 | 54.3 | 30.4 | doesn't fit | doesn't fit | — |
| RTX 5060 Ti 16GB | 16GB | $429 | 76.6 | 41.9 | doesn't fit | doesn't fit | $0.14/hr |
| RTX 5070 Ti | 16GB | $749 | 139 | 79.0 | doesn't fit | doesn't fit | $0.15/hr |
| RTX 5080 | 16GB | $999 | 154 | 87.0 | doesn't fit | doesn't fit | $0.21/hr |
| RTX 3090 | 24GB | $1,499 | 138 | 81.1 | 202 | 38.0 | $0.12/hr |
| RTX 4090 | 24GB | $1,599 | 164 | 96.4 | 260 | 44.3 | $0.34/hr |
| RTX 5090 | 32GB | $1,999 | 244 | 143 | 345 | 71.2 | $0.39/hr |
Our own llama-bench runs: 512-token prompt, 128 generated tokens, three runs after a warmup. 'Doesn't fit' means the model at Q4_K_M plus context exceeds the card's VRAM. Rent is the cheapest hourly rate we tracked on RunPod or Vast.ai in October 2026.
Qwen3 8B: tokens per second on each card
Qwen3 8B: tok/s per $1,000 of launch price
VRAM decides what you can run. A 12GB card like the RTX 3060 runs models up to about 14B at Q4. 16GB adds headroom and longer context, but still stops short of the 30B class. 24GB is where it opens up: Qwen3 32B runs on an RTX 3090 at 38.0 tok/s and on an RTX 4090 at 44.3, and the RTX 5090's 32GB takes it to 71.2. Buy the memory for the biggest model you actually want, then pick the fastest card at that size.
Speed comes from memory bandwidth. Token generation reads the whole model for every token, so the ranking follows bandwidth, not shader count. That's why the RTX 3090 (138 tok/s on Qwen3 8B) stays close to the RTX 4090 (164) for text, while the 4090 is about twice as fast for image generation.
Blackwell moved the 16GB tier. The RTX 5070 Ti runs Qwen3 8B at 139 tok/s, 2.6 times the RTX 4060 Ti 16GB at 54.3, because its memory bus is far wider. Even the RTX 5060 Ti 16GB (76.6 tok/s) beats the 4060 Ti by 41%. If you're buying a 16GB card for AI in 2026, the 50 series is the one to buy.
Mixture-of-experts models are the cheat code for 24GB. Qwen3 30B-A3B only computes about 3B parameters per token, so on a 3090 it runs at 202 tok/s while the dense Qwen3 32B manages 38.0. If your card has 24GB, an MoE model gives you big-model knowledge at small-model speed.
Our picks. On a tight budget, the RTX 3060 12GB is still a real LLM card at 63.8 tok/s on an 8B model. For a new 16GB card, the RTX 5070 Ti is the sweet spot. For 32B models, a used RTX 3090 is the cheapest way into 24GB, and an RTX 4090 adds speed for image and video work. The RTX 5090 is for people who want 32B models fast and can justify $1,999.
Or rent it. At the rates we track, a million tokens of Qwen3 8B costs $0.25 on a rented RTX 3090 and $0.44 on a rented 5090. Unless the card will be busy most of the day, renting is cheaper; our rent-vs-buy guide has the break-even maths.
Buy VRAM first, then bandwidth. 12GB runs models up to about 14B, 24GB opens up 32B models and fast MoE models, and 32GB runs 32B models at 71.2 tok/s. Best value per dollar is the RTX 3060; the best 16GB card is the RTX 5070 Ti (139 tok/s on Qwen3 8B); the cheapest way into 24GB is a used RTX 3090.
All speeds are our own llama.cpp llama-bench measurements at Q4_K_M with full GPU offload (512-token prompt, 128 generated tokens, three runs after a warmup), on rented cards checked for power caps before each run. Launch prices are the manufacturer's US launch MSRP (RTX 5060 Ti and 4060 Ti: the 16GB versions). Rental rates are the cheapest hourly prices we tracked on RunPod and Vast.ai in October 2026.