Serving a language model to many users · 4 models · 10 GPUs measured first-party · Updated October 2026
Which graphics card to use for serving a language model to many users, from first-party measurements of Qwen2.5 1.5B served, SmolLM2 1.7B served, TinyLlama 1.1B served and more on 10 GPUs.

9186.4 serve tok/s on Qwen2.5 1.5B served, the ceiling. Measured on our bench. 192GB of VRAM, $40,000 at launch.

5786.7 serve tok/s on Qwen2.5 1.5B served, lowest launch price that still fits. Measured on our bench. 96GB of VRAM, $8,565 at launch.
Serving means running one model for many people at once: an API, a chatbot behind a website, an agent fleet. A serving engine (vLLM here) batches requests together, so throughput is measured across all users at once, which is a very different number from single-user chat speed.
We measured 4 models for serving a language model to many users on 10 GPUs. Speed is tokens generated per second across all concurrent requests. Every number below is a first-party run on our own harness; cards absent from a model's chart have not been run on it yet.
Qwen2.5 1.5B served: serve tok/s by GPU
SmolLM2 1.7B served: serve tok/s by GPU
TinyLlama 1.1B served: serve tok/s by GPU
Qwen2.5 7B served: serve tok/s by GPU
Which models fit which card, for serving a language model to many users
| Model | VRAM used | 8GB card | 12GB card | 16GB card | 24GB card | 32GB card | Licence |
|---|---|---|---|---|---|---|---|
| Qwen2.5 1.5B served | 12.5GB | No | No | Yes | Yes | Yes | Apache-2.0 |
| SmolLM2 1.7B served | 13.5GB | No | No | Yes | Yes | Yes | Apache-2.0 |
| TinyLlama 1.1B served | 13.5GB | No | No | Yes | Yes | Yes | Apache-2.0 |
| Qwen2.5 7B served | 19.4GB | No | No | No | Yes | Yes | Apache-2.0 |
From the lowest VRAM peak we measured for each model, plus 5% headroom. 'No' means it did not fit in that much memory at our settings, not that no setting ever could.
Every GPU x every serving a language model to many users model (serve tok/s)
| GPU | Qwen2.5 1.5B | SmolLM2 1.7B | TinyLlama 1.1B | Qwen2.5 7B |
|---|---|---|---|---|
| NVIDIA B200 | 9186.4 | 9672.9 | 11562.2 | 5799.1 |
| NVIDIA H200 | 7541.6 | 7107.8 | 9137.7 | 4394.8 |
| NVIDIA H100 80GB HBM3 | 6794.8 | 6509.9 | 8336.7 | 3601.1 |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 5786.7 | 5642.0 | 7678.7 | 2282.8 |
| NVIDIA L40S | 4686.8 | 3525.9 | 6278.4 | 1406.8 |
| NVIDIA A100 80GB SXM4 | 4564.7 | 4351.4 | 5701.5 | 2187.2 |
| NVIDIA A100 40GB SXM4 | 4111.1 | 3865.2 | 3221.3 | 1769.6 |
| NVIDIA A10G | 2702.5 | 2242.7 | 3558.2 | 858.7 |
| NVIDIA L4 | 1827.0 | 1524.9 | 2584.3 | 469.9 |
| NVIDIA T4 | 1337.1 | 1222.5 | 1897.8 | — |
— = not measured on that card yet.
What the numbers show.
Qwen2.5 1.5B served: fastest on the NVIDIA B200 at 9186.4 serve tok/s, 6.87x the slowest card we measured (NVIDIA T4); it used about 12.5GB of VRAM.
SmolLM2 1.7B served: fastest on the NVIDIA B200 at 9672.9 serve tok/s, 7.91x the slowest card we measured (NVIDIA T4); it used about 13.5GB of VRAM.
TinyLlama 1.1B served: fastest on the NVIDIA B200 at 11562.2 serve tok/s, 6.09x the slowest card we measured (NVIDIA T4); it used about 13.5GB of VRAM.
Qwen2.5 7B served: fastest on the NVIDIA B200 at 5799.1 serve tok/s, 12.34x the slowest card we measured (NVIDIA L4); it used about 19.4GB of VRAM.
Which model to pick. On the same card, the NVIDIA B200, TinyLlama 1.1B served runs at 11562.2 serve tok/s in about 13.5GB; SmolLM2 1.7B served runs at 9672.9 serve tok/s in about 13.5GB; Qwen2.5 1.5B served runs at 9186.4 serve tok/s in about 12.5GB; Qwen2.5 7B served runs at 5799.1 serve tok/s in about 19.4GB. TinyLlama 1.1B served gets through the work 2.0x as fast as Qwen2.5 7B served, so the model you choose moves the speed as much as the card does.
How it compares. B200: Qwen2.5 1.5B served 9186.4 serve tok/s, SmolLM2 1.7B served 9672.9, TinyLlama 1.1B served 11562.2, Qwen2.5 7B served 5799.1. 2 of 3 beat Qwen2.5 1.5B served here.
Cost on a rented GPU. 1M served tokens of Qwen2.5 1.5B served: $0.028 on a T4 ($0.14/hr, 12 min), $0.18 on a B200 ($5.98/hr, 2 min, 6.4x the cost).
Qwen2.5 1.5B served: cost per 1M served tokens on rented GPUs
| GPU | Cheapest rate | Speed (serve tok/s) | Cost per 1M served tokens |
|---|---|---|---|
| NVIDIA T4 | $0.14/hr | 1337.1 | $0.028 |
| NVIDIA A100 40GB SXM4 | $0.47/hr | 4111.1 | $0.032 |
| NVIDIA L40S | $0.79/hr | 4686.8 | $0.047 |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | $1.08/hr | 5786.7 | $0.052 |
| NVIDIA A100 80GB SXM4 | $0.95/hr | 4564.7 | $0.058 |
| NVIDIA L4 | $0.44/hr | 1827 | $0.067 |
| NVIDIA H100 80GB HBM3 | $2.14/hr | 6794.8 | $0.087 |
| NVIDIA H200 | $3.59/hr | 7541.6 | $0.13 |
| NVIDIA B200 | $5.98/hr | 9186.4 | $0.18 |
Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.
Speed tiers for Qwen2.5 1.5B served. 5000+ serve tok/s: 4 (B200, H200, H100 80GB HBM3); 1000-5000 serve tok/s: 6 (L40S, A100 80GB SXM4, A100 40GB SXM4). 5,000 serve tok/s is enough for a busy API.
VRAM for Qwen2.5 1.5B served. Measured peak 11.9GB, so 16GB is the smallest common card size; smallest card it ran on: T4 (16GB).
Power on Qwen2.5 1.5B served. Most efficient: H200, 219W, 8.1 Wh per 1M served tokens. Hungriest: B200, 332W, 10.1 Wh.
For serving a language model to many users, the NVIDIA B200 is the fastest card we measured. We have not measured a consumer card on this job yet; the picks above are datacenter and workstation hardware. Check the fit table before buying: VRAM, not speed, is what rules a card out.
Each model runs a fixed workload on every card: a warmup, then timed runs with power, temperature and VRAM sampled every half second through NVML. Models run at the precision and settings from their model card. Datacenter cards run on Modal; consumer cards on rented machines. Non-commercially licensed models are not part of this page.