Serving a language model to many users · 4 models · 10 GPUs measured first-party · Updated October 2026

Best GPU for Serving a language model to many users

Which graphics card to use for serving a language model to many users, from first-party measurements of Qwen2.5 1.5B served, SmolLM2 1.7B served, TinyLlama 1.1B served and more on 10 GPUs.

Fastest we measured
NVIDIA B200

NVIDIA B200

9186.4 serve tok/s on Qwen2.5 1.5B served, the ceiling. Measured on our bench. 192GB of VRAM, $40,000 at launch.

Pros
  • 9186.4 serve tok/s on Qwen2.5 1.5B served
  • 192GB, clears the Qwen2.5 1.5B served floor
  • Rentable by the hour rather than bought
Cons
  • 1000W board rating
  • Datacenter or workstation hardware, not a retail purchase
Cheapest card that runs it
NVIDIA RTX PRO 6000 Blackwell Workstation Edition

NVIDIA RTX PRO 6000 Blackwell Workstation Edition

5786.7 serve tok/s on Qwen2.5 1.5B served, lowest launch price that still fits. Measured on our bench. 96GB of VRAM, $8,565 at launch.

Pros
  • 5786.7 serve tok/s on Qwen2.5 1.5B served
  • 96GB, clears the Qwen2.5 1.5B served floor
  • Rentable by the hour rather than bought
Cons
  • 600W board rating
  • Datacenter or workstation hardware, not a retail purchase
4
Models measured
Qwen2.5 1.5B served, SmolLM2 1.7B served, TinyLlama 1.1B served and more
10
GPUs measured
first-party runs, not spec-sheet estimates
11562.2serve tok/s
Fastest: NVIDIA B200
on TinyLlama 1.1B served
12.5GB
Lightest model's VRAM need
measured peak, +5% headroom

Serving means running one model for many people at once: an API, a chatbot behind a website, an agent fleet. A serving engine (vLLM here) batches requests together, so throughput is measured across all users at once, which is a very different number from single-user chat speed.

We measured 4 models for serving a language model to many users on 10 GPUs. Speed is tokens generated per second across all concurrent requests. Every number below is a first-party run on our own harness; cards absent from a model's chart have not been run on it yet.

Qwen2.5 1.5B served: serve tok/s by GPU

NVIDIA B200
9186.4 serve tok/s
NVIDIA H200
7541.6 serve tok/s
NVIDIA H100 80GB HBM3
6794.8 serve tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
5786.7 serve tok/s
NVIDIA L40S
4686.8 serve tok/s
NVIDIA A100 80GB SXM4
4564.7 serve tok/s
NVIDIA A100 40GB SXM4
4111.1 serve tok/s
NVIDIA A10G
2702.5 serve tok/s
NVIDIA L4
1827 serve tok/s
NVIDIA T4
1337.1 serve tok/s

SmolLM2 1.7B served: serve tok/s by GPU

NVIDIA B200
9672.9 serve tok/s
NVIDIA H200
7107.8 serve tok/s
NVIDIA H100 80GB HBM3
6509.9 serve tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
5642 serve tok/s
NVIDIA A100 80GB SXM4
4351.4 serve tok/s
NVIDIA A100 40GB SXM4
3865.2 serve tok/s
NVIDIA L40S
3525.9 serve tok/s
NVIDIA A10G
2242.7 serve tok/s
NVIDIA L4
1524.9 serve tok/s
NVIDIA T4
1222.5 serve tok/s

TinyLlama 1.1B served: serve tok/s by GPU

NVIDIA B200
11562.2 serve tok/s
NVIDIA H200
9137.7 serve tok/s
NVIDIA H100 80GB HBM3
8336.7 serve tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
7678.7 serve tok/s
NVIDIA L40S
6278.4 serve tok/s
NVIDIA A100 80GB SXM4
5701.5 serve tok/s
NVIDIA A10G
3558.2 serve tok/s
NVIDIA A100 40GB SXM4
3221.3 serve tok/s
NVIDIA L4
2584.3 serve tok/s
NVIDIA T4
1897.8 serve tok/s

Qwen2.5 7B served: serve tok/s by GPU

NVIDIA B200
5799.1 serve tok/s
NVIDIA H200
4394.8 serve tok/s
NVIDIA H100 80GB HBM3
3601.1 serve tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
2282.8 serve tok/s
NVIDIA A100 80GB SXM4
2187.2 serve tok/s
NVIDIA A100 40GB SXM4
1769.6 serve tok/s
NVIDIA L40S
1406.8 serve tok/s
NVIDIA A10G
858.7 serve tok/s
NVIDIA L4
469.9 serve tok/s

Which models fit which card, for serving a language model to many users

Qwen2.5 1.5B served12.5GB
SmolLM2 1.7B served13.5GB
TinyLlama 1.1B served13.5GB
Qwen2.5 7B served19.4GB
ModelVRAM used8GB card12GB card16GB card24GB card32GB cardLicence
Qwen2.5 1.5B served12.5GBNoNoYesYesYesApache-2.0
SmolLM2 1.7B served13.5GBNoNoYesYesYesApache-2.0
TinyLlama 1.1B served13.5GBNoNoYesYesYesApache-2.0
Qwen2.5 7B served19.4GBNoNoNoYesYesApache-2.0

From the lowest VRAM peak we measured for each model, plus 5% headroom. 'No' means it did not fit in that much memory at our settings, not that no setting ever could.

Every GPU x every serving a language model to many users model (serve tok/s)

NVIDIA B2009186.4
NVIDIA H2007541.6
NVIDIA H100 80GB HBM36794.8
NVIDIA RTX PRO 6000 Blackwell Workstation Edition5786.7
NVIDIA L40S4686.8
NVIDIA A100 80GB SXM44564.7
NVIDIA A100 40GB SXM44111.1
NVIDIA A10G2702.5
NVIDIA L41827.0
NVIDIA T41337.1
GPUQwen2.5 1.5BSmolLM2 1.7BTinyLlama 1.1BQwen2.5 7B
NVIDIA B2009186.49672.911562.25799.1
NVIDIA H2007541.67107.89137.74394.8
NVIDIA H100 80GB HBM36794.86509.98336.73601.1
NVIDIA RTX PRO 6000 Blackwell Workstation Edition5786.75642.07678.72282.8
NVIDIA L40S4686.83525.96278.41406.8
NVIDIA A100 80GB SXM44564.74351.45701.52187.2
NVIDIA A100 40GB SXM44111.13865.23221.31769.6
NVIDIA A10G2702.52242.73558.2858.7
NVIDIA L41827.01524.92584.3469.9
NVIDIA T41337.11222.51897.8—

— = not measured on that card yet.

What the numbers show.

Qwen2.5 1.5B served: fastest on the NVIDIA B200 at 9186.4 serve tok/s, 6.87x the slowest card we measured (NVIDIA T4); it used about 12.5GB of VRAM.

SmolLM2 1.7B served: fastest on the NVIDIA B200 at 9672.9 serve tok/s, 7.91x the slowest card we measured (NVIDIA T4); it used about 13.5GB of VRAM.

TinyLlama 1.1B served: fastest on the NVIDIA B200 at 11562.2 serve tok/s, 6.09x the slowest card we measured (NVIDIA T4); it used about 13.5GB of VRAM.

Qwen2.5 7B served: fastest on the NVIDIA B200 at 5799.1 serve tok/s, 12.34x the slowest card we measured (NVIDIA L4); it used about 19.4GB of VRAM.

Which model to pick. On the same card, the NVIDIA B200, TinyLlama 1.1B served runs at 11562.2 serve tok/s in about 13.5GB; SmolLM2 1.7B served runs at 9672.9 serve tok/s in about 13.5GB; Qwen2.5 1.5B served runs at 9186.4 serve tok/s in about 12.5GB; Qwen2.5 7B served runs at 5799.1 serve tok/s in about 19.4GB. TinyLlama 1.1B served gets through the work 2.0x as fast as Qwen2.5 7B served, so the model you choose moves the speed as much as the card does.

How it compares. B200: Qwen2.5 1.5B served 9186.4 serve tok/s, SmolLM2 1.7B served 9672.9, TinyLlama 1.1B served 11562.2, Qwen2.5 7B served 5799.1. 2 of 3 beat Qwen2.5 1.5B served here.

Cost on a rented GPU. 1M served tokens of Qwen2.5 1.5B served: $0.028 on a T4 ($0.14/hr, 12 min), $0.18 on a B200 ($5.98/hr, 2 min, 6.4x the cost).

Qwen2.5 1.5B served: cost per 1M served tokens on rented GPUs

NVIDIA T4$0.14/hr
NVIDIA A100 40GB SXM4$0.47/hr
NVIDIA L40S$0.79/hr
NVIDIA RTX PRO 6000 Blackwell Workstation Edition$1.08/hr
NVIDIA A100 80GB SXM4$0.95/hr
NVIDIA L4$0.44/hr
NVIDIA H100 80GB HBM3$2.14/hr
NVIDIA H200$3.59/hr
NVIDIA B200$5.98/hr
GPUCheapest rateSpeed (serve tok/s)Cost per 1M served tokens
NVIDIA T4$0.14/hr1337.1$0.028
NVIDIA A100 40GB SXM4$0.47/hr4111.1$0.032
NVIDIA L40S$0.79/hr4686.8$0.047
NVIDIA RTX PRO 6000 Blackwell Workstation Edition$1.08/hr5786.7$0.052
NVIDIA A100 80GB SXM4$0.95/hr4564.7$0.058
NVIDIA L4$0.44/hr1827$0.067
NVIDIA H100 80GB HBM3$2.14/hr6794.8$0.087
NVIDIA H200$3.59/hr7541.6$0.13
NVIDIA B200$5.98/hr9186.4$0.18

Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.

Speed tiers for Qwen2.5 1.5B served. 5000+ serve tok/s: 4 (B200, H200, H100 80GB HBM3); 1000-5000 serve tok/s: 6 (L40S, A100 80GB SXM4, A100 40GB SXM4). 5,000 serve tok/s is enough for a busy API.

VRAM for Qwen2.5 1.5B served. Measured peak 11.9GB, so 16GB is the smallest common card size; smallest card it ran on: T4 (16GB).

Power on Qwen2.5 1.5B served. Most efficient: H200, 219W, 8.1 Wh per 1M served tokens. Hungriest: B200, 332W, 10.1 Wh.

Our verdict

For serving a language model to many users, the NVIDIA B200 is the fastest card we measured. We have not measured a consumer card on this job yet; the picks above are datacenter and workstation hardware. Check the fit table before buying: VRAM, not speed, is what rules a card out.

FAQ

What is the fastest GPU for serving a language model to many users?
In our runs, the NVIDIA B200 at 11562.2 serve tok/s on TinyLlama 1.1B served. We measured 4 models on 10 GPUs for this page.
How much VRAM do I need for serving a language model to many users?
The lightest model here, Qwen2.5 1.5B served, used about 12.5GB. The table above shows which models fit 8, 12, 16, 24 and 32GB cards, from measured peaks.
Are these numbers measured or estimated?
Measured. Every number on this page is a first-party run on our own harness, with power and VRAM sampled during the run. Cards we have not run yet are simply absent, not filled in.
What GPU do I need to run Qwen2.5 1.5B served?
About 13GB. Smallest card that ran it: NVIDIA T4 (16GB).
How much does it cost to run Qwen2.5 1.5B served in the cloud?
$0.028 per 1M served tokens on a NVIDIA T4 at $0.14/hr, cheapest of 9 rentable cards we measured.
Can I run Qwen2.5 1.5B served on a 12GB, 16GB or 24GB card?
It used 11.9GB at the precision we tested. 12GB: no; 16GB: yes; 24GB: yes.
Is the H100 80GB HBM3 or the A100 80GB SXM4 faster for Qwen2.5 1.5B served?
The H100 80GB HBM3: 6794.8 vs 4564.7 serve tok/s, 49% faster on our bench.

How we test

Each model runs a fixed workload on every card: a warmup, then timed runs with power, temperature and VRAM sampled every half second through NVML. Models run at the precision and settings from their model card. Datacenter cards run on Modal; consumer cards on rented machines. Non-commercially licensed models are not part of this page.