Speculative decoding · 1 model · 6 GPUs measured first-party · Updated October 2026

Best GPU for Speculative decoding

Which graphics card to use for speculative decoding, from first-party measurements of Qwen2.5 1.5B + 0.5B draft on 6 GPUs.

Fastest we measured
NVIDIA L4

NVIDIA L4

0.97 x vs solo on Qwen2.5 1.5B + 0.5B draft, the ceiling. Measured on our bench. 24GB of VRAM, $2,500 at launch.

Pros
  • 0.97 x vs solo on Qwen2.5 1.5B + 0.5B draft
  • 24GB, clears the Qwen2.5 1.5B + 0.5B draft floor
  • Rentable by the hour rather than bought
Cons
  • 72W board rating
  • Datacenter or workstation hardware, not a retail purchase
1
Models measured
Qwen2.5 1.5B + 0.5B draft
6
GPUs measured
first-party runs, not spec-sheet estimates
0.97x vs solo
Fastest: NVIDIA L4
on Qwen2.5 1.5B + 0.5B draft
12GB
Lightest model's VRAM need
measured peak, +5% headroom

Speculative decoding speeds up a language model by letting a tiny draft model guess several tokens ahead and having the big model check them in one pass. Whether it helps depends on the card: on some GPUs the check is nearly free, on others the draft model just adds work.

We measured 1 model for speculative decoding on 6 GPUs. Speed is speed relative to running the main model alone (1.5x = 50% faster). Every number below is a first-party run on our own harness; cards absent from a model's chart have not been run on it yet.

Qwen2.5 1.5B + 0.5B draft: x vs solo by GPU

NVIDIA L4
0.97 x vs solo
NVIDIA L40S
0.95 x vs solo
NVIDIA A10G
0.92 x vs solo
NVIDIA T4
0.9 x vs solo
NVIDIA H100 80GB HBM3
0.72 x vs solo
NVIDIA A100 40GB SXM4
0.69 x vs solo

Which models fit which card, for speculative decoding

ModelVRAM used8GB card12GB card16GB card24GB card32GB cardLicence
Qwen2.5 1.5B + 0.5B draft12GBNoNoYesYesYesApache-2.0

From the lowest VRAM peak we measured for each model, plus 5% headroom. 'No' means it did not fit in that much memory at our settings, not that no setting ever could.

Qwen2.5 1.5B + 0.5B draft: every GPU we measured

NVIDIA L40.97
NVIDIA L40S0.95
NVIDIA A10G0.92
NVIDIA T40.9
NVIDIA H100 80GB HBM30.72
NVIDIA A100 40GB SXM40.69
GPUx vs soloVRAMPower
NVIDIA L40.9724GB—
NVIDIA L40S0.9548GB—
NVIDIA A10G0.9224GB—
NVIDIA T40.916GB—
NVIDIA H100 80GB HBM30.7280GB—
NVIDIA A100 40GB SXM40.6940GB—

What the numbers show.

Qwen2.5 1.5B + 0.5B draft: fastest on the NVIDIA L4 at 0.97 x vs solo, 1.39x the slowest card we measured (NVIDIA A100 40GB SXM4); it used about 12GB of VRAM.

Does it help on your card? On 6 of the 6 cards we measured, adding the draft model made generation slower, down to 0.69x on the NVIDIA A100 40GB SXM4. It did not help on any card we measured. Speculative decoding pays off when the big model is slow enough that guessing ahead is cheaper than waiting, and when the draft guesses right most of the time. A small target model on a fast card is already quick, so the draft's extra work and its rejected guesses cost more than they save. Try it on a large model on a card that is short on bandwidth, and measure before you leave it on.

Our verdict

For speculative decoding, the NVIDIA L4 is the fastest card we measured. We have not measured a consumer card on this job yet; the picks above are datacenter and workstation hardware. Check the fit table before buying: VRAM, not speed, is what rules a card out.

FAQ

What is the fastest GPU for speculative decoding?
In our runs, the NVIDIA L4 at 0.97 x vs solo on Qwen2.5 1.5B + 0.5B draft. We measured 1 model on 6 GPUs for this page.
How much VRAM do I need for speculative decoding?
The lightest model here, Qwen2.5 1.5B + 0.5B draft, used about 12GB. The table above shows which models fit 8, 12, 16, 24 and 32GB cards, from measured peaks.
Are these numbers measured or estimated?
Measured. Every number on this page is a first-party run on our own harness, with power and VRAM sampled during the run. Cards we have not run yet are simply absent, not filled in.
Does speculative decoding always make a model faster?
No. With Qwen2.5 1.5B + 0.5B draft it was slower on 6 of the 6 cards we measured; the best result was 0.97x on the NVIDIA L4. It is a setting to test on your own card and model, not a free speedup.

How we test

Each model runs a fixed workload on every card: a warmup, then timed runs with power, temperature and VRAM sampled every half second through NVML. Models run at the precision and settings from their model card. Datacenter cards run on Modal; consumer cards on rented machines. Non-commercially licensed models are not part of this page.