Buyer's guide · local LLMs · 8 consumer cards measured · Updated October 2026

Best GPU for Running LLMs Locally

We ran the same language models on eight consumer cards, from the 12GB RTX 3060 to the 32GB RTX 5090. Two numbers decide which card is right for you: how much VRAM it has, which decides which models load at all, and how fast it generates once they do. Here's both, measured.

63.8tok/s
Qwen3 8B on an RTX 3060
the cheapest card here, still faster than you read
244tok/s
Qwen3 8B on an RTX 5090
the fastest consumer card
2.6x
RTX 5070 Ti vs RTX 4060 Ti
same 16GB, one generation apart
24GB
Where 32B models start
Qwen3 32B needs about 25GB with context

Generation speed on the same models (tok/s, Q4_K_M, llama.cpp)

RTX 306012GB
RTX 4060 Ti 16GB16GB
RTX 5060 Ti 16GB16GB
RTX 5070 Ti16GB
RTX 508016GB
RTX 309024GB
RTX 409024GB
RTX 509032GB
GPUVRAMLaunch priceQwen3 8BQwen3 14BQwen3 30B-A3BQwen3 32BCheapest rent
RTX 306012GB$32963.836.5doesn't fitdoesn't fit$0.04/hr
RTX 4060 Ti 16GB16GB$49954.330.4doesn't fitdoesn't fit—
RTX 5060 Ti 16GB16GB$42976.641.9doesn't fitdoesn't fit$0.14/hr
RTX 5070 Ti16GB$74913979.0doesn't fitdoesn't fit$0.15/hr
RTX 508016GB$99915487.0doesn't fitdoesn't fit$0.21/hr
RTX 309024GB$1,49913881.120238.0$0.12/hr
RTX 409024GB$1,59916496.426044.3$0.34/hr
RTX 509032GB$1,99924414334571.2$0.39/hr

Our own llama-bench runs: 512-token prompt, 128 generated tokens, three runs after a warmup. 'Doesn't fit' means the model at Q4_K_M plus context exceeds the card's VRAM. Rent is the cheapest hourly rate we tracked on RunPod or Vast.ai in October 2026.

Qwen3 8B: tokens per second on each card

RTX 3060
63.8 tok/s
RTX 4060 Ti 16GB
54.3 tok/s
RTX 5060 Ti 16GB
76.6 tok/s
RTX 5070 Ti
139.3 tok/s
RTX 5080
153.9 tok/s
RTX 3090
137.6 tok/s
RTX 4090
164.3 tok/s
RTX 5090
243.9 tok/s

Qwen3 8B: tok/s per $1,000 of launch price

RTX 3060
194 tok/s per $1k
RTX 4060 Ti 16GB
108.9 tok/s per $1k
RTX 5060 Ti 16GB
178.6 tok/s per $1k
RTX 5070 Ti
185.9 tok/s per $1k
RTX 5080
154 tok/s per $1k
RTX 3090
91.8 tok/s per $1k
RTX 4090
102.8 tok/s per $1k
RTX 5090
122 tok/s per $1k

VRAM decides what you can run. A 12GB card like the RTX 3060 runs models up to about 14B at Q4. 16GB adds headroom and longer context, but still stops short of the 30B class. 24GB is where it opens up: Qwen3 32B runs on an RTX 3090 at 38.0 tok/s and on an RTX 4090 at 44.3, and the RTX 5090's 32GB takes it to 71.2. Buy the memory for the biggest model you actually want, then pick the fastest card at that size.

Speed comes from memory bandwidth. Token generation reads the whole model for every token, so the ranking follows bandwidth, not shader count. That's why the RTX 3090 (138 tok/s on Qwen3 8B) stays close to the RTX 4090 (164) for text, while the 4090 is about twice as fast for image generation.

Blackwell moved the 16GB tier. The RTX 5070 Ti runs Qwen3 8B at 139 tok/s, 2.6 times the RTX 4060 Ti 16GB at 54.3, because its memory bus is far wider. Even the RTX 5060 Ti 16GB (76.6 tok/s) beats the 4060 Ti by 41%. If you're buying a 16GB card for AI in 2026, the 50 series is the one to buy.

Mixture-of-experts models are the cheat code for 24GB. Qwen3 30B-A3B only computes about 3B parameters per token, so on a 3090 it runs at 202 tok/s while the dense Qwen3 32B manages 38.0. If your card has 24GB, an MoE model gives you big-model knowledge at small-model speed.

Our picks. On a tight budget, the RTX 3060 12GB is still a real LLM card at 63.8 tok/s on an 8B model. For a new 16GB card, the RTX 5070 Ti is the sweet spot. For 32B models, a used RTX 3090 is the cheapest way into 24GB, and an RTX 4090 adds speed for image and video work. The RTX 5090 is for people who want 32B models fast and can justify $1,999.

Or rent it. At the rates we track, a million tokens of Qwen3 8B costs $0.25 on a rented RTX 3090 and $0.44 on a rented 5090. Unless the card will be busy most of the day, renting is cheaper; our rent-vs-buy guide has the break-even maths.

Our verdict

Buy VRAM first, then bandwidth. 12GB runs models up to about 14B, 24GB opens up 32B models and fast MoE models, and 32GB runs 32B models at 71.2 tok/s. Best value per dollar is the RTX 3060; the best 16GB card is the RTX 5070 Ti (139 tok/s on Qwen3 8B); the cheapest way into 24GB is a used RTX 3090.

FAQ

What is the best GPU for running LLMs locally?
For most people, a 16GB RTX 5070 Ti (139 tok/s on Qwen3 8B) or, for 32B models, a 24GB RTX 3090 or 4090. The RTX 5090 is fastest at 244 tok/s and the only consumer card with 32GB.
How much VRAM do I need to run a local LLM?
About 6GB for an 8B model at Q4, 11GB for 14B, and around 25GB for a dense 32B model with context. 12GB covers small and mid models; 24GB is where 32B models fit.
Is the RTX 3090 still good for LLMs?
Yes. For text generation it runs Qwen3 8B at 138 tok/s, about 84% of an RTX 4090, because the two have similar memory bandwidth. It falls well behind for image and video.
Is the RTX 3060 good enough for local AI?
For models up to about 14B, yes: Qwen3 8B runs at 63.8 tok/s, faster than you read. Its 12GB also beats the 8GB RTX 4060 for AI, even though the 4060 is newer.

How we test

All speeds are our own llama.cpp llama-bench measurements at Q4_K_M with full GPU offload (512-token prompt, 128 generated tokens, three runs after a warmup), on rented cards checked for power caps before each run. Launch prices are the manufacturer's US launch MSRP (RTX 5060 Ti and 4060 Ti: the 16GB versions). Rental rates are the cheapest hourly prices we tracked on RunPod and Vast.ai in October 2026.