Speculative decoding · 1 model · 6 GPUs measured first-party · Updated October 2026
Which graphics card to use for speculative decoding, from first-party measurements of Qwen2.5 1.5B + 0.5B draft on 6 GPUs.

0.97 x vs solo on Qwen2.5 1.5B + 0.5B draft, the ceiling. Measured on our bench. 24GB of VRAM, $2,500 at launch.
Speculative decoding speeds up a language model by letting a tiny draft model guess several tokens ahead and having the big model check them in one pass. Whether it helps depends on the card: on some GPUs the check is nearly free, on others the draft model just adds work.
We measured 1 model for speculative decoding on 6 GPUs. Speed is speed relative to running the main model alone (1.5x = 50% faster). Every number below is a first-party run on our own harness; cards absent from a model's chart have not been run on it yet.
Qwen2.5 1.5B + 0.5B draft: x vs solo by GPU
Which models fit which card, for speculative decoding
| Model | VRAM used | 8GB card | 12GB card | 16GB card | 24GB card | 32GB card | Licence |
|---|---|---|---|---|---|---|---|
| Qwen2.5 1.5B + 0.5B draft | 12GB | No | No | Yes | Yes | Yes | Apache-2.0 |
From the lowest VRAM peak we measured for each model, plus 5% headroom. 'No' means it did not fit in that much memory at our settings, not that no setting ever could.
Qwen2.5 1.5B + 0.5B draft: every GPU we measured
| GPU | x vs solo | VRAM | Power |
|---|---|---|---|
| NVIDIA L4 | 0.97 | 24GB | — |
| NVIDIA L40S | 0.95 | 48GB | — |
| NVIDIA A10G | 0.92 | 24GB | — |
| NVIDIA T4 | 0.9 | 16GB | — |
| NVIDIA H100 80GB HBM3 | 0.72 | 80GB | — |
| NVIDIA A100 40GB SXM4 | 0.69 | 40GB | — |
What the numbers show.
Qwen2.5 1.5B + 0.5B draft: fastest on the NVIDIA L4 at 0.97 x vs solo, 1.39x the slowest card we measured (NVIDIA A100 40GB SXM4); it used about 12GB of VRAM.
Does it help on your card? On 6 of the 6 cards we measured, adding the draft model made generation slower, down to 0.69x on the NVIDIA A100 40GB SXM4. It did not help on any card we measured. Speculative decoding pays off when the big model is slow enough that guessing ahead is cheaper than waiting, and when the draft guesses right most of the time. A small target model on a fast card is already quick, so the draft's extra work and its rejected guesses cost more than they save. Try it on a large model on a card that is short on bandwidth, and measure before you leave it on.
For speculative decoding, the NVIDIA L4 is the fastest card we measured. We have not measured a consumer card on this job yet; the picks above are datacenter and workstation hardware. Check the fit table before buying: VRAM, not speed, is what rules a card out.
Each model runs a fixed workload on every card: a warmup, then timed runs with power, temperature and VRAM sampled every half second through NVML. Models run at the precision and settings from their model card. Datacenter cards run on Modal; consumer cards on rented machines. Non-commercially licensed models are not part of this page.