Can You Run TinyLlama 1.1B served? Every GPU, Tested

A model that doesn't fit doesn't run slowly, it doesn't run. Here's every card from our bench, sorted by measured speed.

Weights: TinyLlama/TinyLlama-1.1B-Chat-v1.0

GPUs that run TinyLlama 1.1B served (10)

GPUSpeedVRAMBasis
NVIDIA B20011562.2 serve tok/s192GB✓ Measured
NVIDIA H2009137.7 serve tok/s141GB✓ Measured
NVIDIA H100 80GB HBM38336.7 serve tok/s80GB✓ Measured
NVIDIA RTX PRO 6000 Blackwell Workstation Edition7678.7 serve tok/s96GB✓ Measured
NVIDIA L40S6278.4 serve tok/s48GB✓ Measured
NVIDIA A100 80GB SXM45701.5 serve tok/s80GB✓ Measured
NVIDIA A10G3558.2 serve tok/s24GB✓ Measured
NVIDIA A100 40GB SXM43221.3 serve tok/s40GBEst.
NVIDIA L42584.3 serve tok/s24GB✓ Measured
NVIDIA T41897.8 serve tok/s16GB✓ Measured

Quick answers

What is the fastest GPU for TinyLlama 1.1B served?

NVIDIA B200 is the fastest card we've measured running TinyLlama 1.1B served, at 11562.2 serve tok/s.

What is the cheapest GPU that can run TinyLlama 1.1B served?

NVIDIA T4 ($2,299 MSRP) is the cheapest card on our bench that runs TinyLlama 1.1B served, at 1897.8 serve tok/s.

← All models · AI GPU rankings · How we benchmark