Explainer · small models · with video · Updated October 2026
Bigger isn't automatically better anymore. Small models trained for one job are matching giant general models, and they run faster, fit cheaper GPUs and cost a fraction per token. Here's the evidence, with our own measurements across one model family from 4B to 235B.

Watch: Small AI Models Are Beating the Giants
A 7B model beat o1-preview at math. In early 2025 Microsoft Research published rStar-Math. They took a 7B math model that scored 58.8% on the MATH benchmark and, by teaching it to think longer and check its own steps, pushed it to 90%. OpenAI's o1-preview scored 85.5% on the same test. Same small model, different training, and it beat one of the biggest reasoning models of its time.
That's why I think the future isn't one giant model that knows everything. Giant general models carry a lot of knowledge that has nothing to do with your task. A small model trained for one job, or several specialized models with a bigger one orchestrating and checking their work, can do better at a fraction of the cost.
One model family, six sizes: what each needs and costs
| Model | Size | VRAM at Q4 | Speed on an RTX 5090 (tok/s) | Cheapest per 1M tokens |
|---|---|---|---|---|
| Qwen3 4B | 4B | 4GB | 375.6 | $0.10 on RTX 3060 |
| Qwen3 8B | 8B | 6GB | 243.91 | $0.16 on RTX 3060 |
| Qwen3 14B | 14B | 11GB | 143.11 | $0.27 on RTX 3060 |
| Qwen3 30B-A3B (MoE) | 30B, 3B active | 20GB | 344.67 | $0.31 on RTX 5090 |
| Qwen3 32B | 32B | 25GB | 71.15 | $1.52 on RTX 5090 |
| Qwen3 235B-A22B (MoE) | 235B, 22B active | 160GB | doesn't fit | $21.28 on B300 |
Speeds are our own single-stream llama.cpp measurements at Q4_K_M. Cost per 1M tokens is the cheapest hourly rental we track for any card we measured the model on, divided by its speed. The 235B figure is from an NVIDIA B300; it needs about 160GB and fits no consumer card.
Qwen3 on an RTX 5090: generation speed by model size
Same card, same quantization. The MoE model only computes 3B parameters per token, which is why it runs like a small model.
Smaller means faster, cheaper and easier to fit. On the same RTX 5090, Qwen3 4B generates 375.6 tokens a second and Qwen3 32B manages 71.15. The 4B needs about 4GB at Q4 and runs on almost any card; the 32B needs about 25GB. Per million tokens, the 4B costs $0.10 on a rented RTX 3060, while the 235B costs $21.28 on a B300, about 217 times as much.
There's a cost people forget with the giants: loading. A 160GB model takes a long time to load into VRAM, and on per-second billing you pay for every one of those seconds before it writes a single token. Small models load fast, and you can fit several of them on one GPU.
Three ways a small model beats a giant. First, it's trained for one job: a model trained on your task beats a model trained on everything. Second, distillation: a small model learns the reasoning of a bigger one, which is how a lot of the strongest open models were built. Third, it thinks longer: instead of being bigger, it tries many answers and keeps the best one, which is exactly what rStar-Math does.
When the giants still win. Broad knowledge on anything, brand-new tasks nobody trained for, and work that needs a lot of context and reasoning at once, like complex code across a large codebase. That's where a big model earns its cost, and where it works well as the orchestrator over smaller specialists.
For most people right now, a subscription to one of the big chat models is still subsidized and a good deal. For batch work, data jobs and agents doing one thing over and over, a small specialized model on a cheap rented GPU is where the cost per result drops off a cliff.
Small models are winning on cost and speed, and with the right training they can match or beat giants on specific tasks: a 7B model with rStar-Math scored 90% on MATH against o1-preview's 85.5%. On our bench Qwen3 4B costs $0.10 per million tokens on a rented GPU and Qwen3 235B costs $21.28. Use a small specialized model for one repeated job, and keep the giants for broad knowledge, new tasks and complex reasoning.
Speeds are our own llama.cpp llama-bench measurements at Q4_K_M (single stream, 128 generated tokens, three runs after a warmup). VRAM at Q4 comes from our quantization ladder. Cost per 1M tokens is the cheapest hourly rate we tracked on RunPod or Vast.ai on October 5, 2026, for any card we measured the model on, divided by its speed. The rStar-Math results are from Microsoft Research's January 2025 paper. Written from the GPU Battle video 'Small AI Models Are Beating the Giants'.