Fastest GPU For Speculative Decoding

Pairing a small draft model with a large one is sold as a free speedup. Measured per card, it is often a slowdown. Above 1.0x it helped, below 1.0x it cost you.

Qwen2.5 1.5B + 0.5B draft

NVIDIA L40.968
NVIDIA L40S0.946
NVIDIA A10G0.919
NVIDIA T40.9
NVIDIA H100 80GB HBM30.725
NVIDIA A100 40GB SXM40.694

Measured in x vs solo. Longer is faster.

#GPUx vs solo
1NVIDIA L40.97
2NVIDIA L40S0.95
3NVIDIA A10G0.92
4NVIDIA T40.9
5NVIDIA H100 80GB HBM30.73
6NVIDIA A100 40GB SXM40.69

Every number is a first-party run on our own bench. Cards missing from a board have not been measured on that model yet, or cannot fit it.