Text-to-speech · 1 model · 8 GPUs measured first-party · Updated October 2026

Best GPU for Text-to-speech

Which graphics card to use for text-to-speech, from first-party measurements of Kokoro TTS 82M on 8 GPUs.

Fastest we measured
NVIDIA L40S

NVIDIA L40S

244.0 x realtime on Kokoro TTS 82M, the ceiling. Measured on our bench. 48GB of VRAM, $7,500 at launch.

Pros
  • 244.0 x realtime on Kokoro TTS 82M
  • 48GB, clears the Kokoro TTS 82M floor
  • Rentable by the hour rather than bought
Cons
  • 350W board rating
  • Datacenter or workstation hardware, not a retail purchase
1
Models measured
Kokoro TTS 82M
8
GPUs measured
first-party runs, not spec-sheet estimates
244.0x realtime
Fastest: NVIDIA L40S
on Kokoro TTS 82M
1.5GB
Lightest model's VRAM need
measured peak, +5% headroom

Text-to-speech turns written text into a spoken voice: narration for videos, audiobooks, voice agents. Kokoro is a small, high-quality open model, which makes it a good measure of how well a card handles lightweight, latency-sensitive audio work.

We measured 1 model for text-to-speech on 8 GPUs. Speed is times faster than realtime. Every number below is a first-party run on our own harness; cards absent from a model's chart have not been run on it yet.

Kokoro TTS 82M: x realtime by GPU

NVIDIA L40S
244 x realtime
NVIDIA H100 80GB HBM3
222 x realtime
NVIDIA H200
170.8 x realtime
NVIDIA A100 40GB SXM4
142 x realtime
NVIDIA A100 80GB SXM4
110.2 x realtime
NVIDIA A10G
101.1 x realtime
NVIDIA L4
97.35 x realtime
NVIDIA T4
42.88 x realtime

Which models fit which card, for text-to-speech

ModelVRAM used8GB card12GB card16GB card24GB card32GB cardLicence
Kokoro TTS 82M1.5GBYesYesYesYesYesApache-2.0

From the lowest VRAM peak we measured for each model, plus 5% headroom. 'No' means it did not fit in that much memory at our settings, not that no setting ever could.

Kokoro TTS 82M: every GPU we measured

NVIDIA L40S244
NVIDIA H100 80GB HBM3222
NVIDIA H200170.8
NVIDIA A100 40GB SXM4142
NVIDIA A100 80GB SXM4110.2
NVIDIA A10G101.1
NVIDIA L497.35
NVIDIA T442.88
GPUx realtimeVRAMPower
NVIDIA L40S24448GB83.2 W
NVIDIA H100 80GB HBM322280GB125.5 W
NVIDIA H200170.8141GB120.3 W
NVIDIA A100 40GB SXM414240GB60.1 W
NVIDIA A100 80GB SXM4110.280GB72.3 W
NVIDIA A10G101.124GB65.5 W
NVIDIA L497.3524GB33.6 W
NVIDIA T442.8816GB51.0 W

What the numbers show.

Kokoro TTS 82M: fastest on the NVIDIA L40S at 244.0 x realtime, 5.69x the slowest card we measured (NVIDIA T4); it used about 1.5GB of VRAM.

How it compares. L40S: Kokoro TTS 82M 244.0 x realtime, Whisper large-v3 193.2, ACE-Step 1.5 20.44, ACE-Step v1 3.5B 13.78, DiffRhythm 2 5.43. Kokoro TTS 82M beats all 4 here.

Cost on a rented GPU. 1 hour of audio of Kokoro TTS 82M: $0.003 on a T4 ($0.14/hr, 1 min), $0.003 on a L40S ($0.79/hr, 0 min, 1.0x the cost).

Kokoro TTS 82M: cost per 1 hour of audio on rented GPUs

NVIDIA T4$0.14/hr
NVIDIA L40S$0.79/hr
NVIDIA A100 40GB SXM4$0.47/hr
NVIDIA L4$0.44/hr
NVIDIA A100 80GB SXM4$0.95/hr
NVIDIA H100 80GB HBM3$2.14/hr
NVIDIA H200$3.59/hr
GPUCheapest rateSpeed (x realtime)Cost per 1 hour of audio
NVIDIA T4$0.14/hr42.88$0.003
NVIDIA L40S$0.79/hr244$0.003
NVIDIA A100 40GB SXM4$0.47/hr142$0.003
NVIDIA L4$0.44/hr97.35$0.005
NVIDIA A100 80GB SXM4$0.95/hr110.2$0.009
NVIDIA H100 80GB HBM3$2.14/hr222$0.010
NVIDIA H200$3.59/hr170.8$0.021

Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.

Speed tiers for Kokoro TTS 82M. 100+ x realtime: 6 (L40S, H100 80GB HBM3, H200); 10-100 x realtime: 2 (L4, T4). 1x is the speed of playback.

VRAM for Kokoro TTS 82M. Measured peak 1.4GB, so 8GB is the smallest common card size; smallest card it ran on: T4 (16GB).

Power on Kokoro TTS 82M. Most efficient: L40S, 83W, 0.3 Wh per 1 hour of audio.

Our verdict

For text-to-speech, the NVIDIA L40S is the fastest card we measured. We have not measured a consumer card on this job yet; the picks above are datacenter and workstation hardware. Check the fit table before buying: VRAM, not speed, is what rules a card out.

FAQ

What is the fastest GPU for text-to-speech?
In our runs, the NVIDIA L40S at 244.0 x realtime on Kokoro TTS 82M. We measured 1 model on 8 GPUs for this page.
How much VRAM do I need for text-to-speech?
The lightest model here, Kokoro TTS 82M, used about 1.5GB. The table above shows which models fit 8, 12, 16, 24 and 32GB cards, from measured peaks.
Are these numbers measured or estimated?
Measured. Every number on this page is a first-party run on our own harness, with power and VRAM sampled during the run. Cards we have not run yet are simply absent, not filled in.
What GPU do I need to run Kokoro TTS 82M?
About 2GB. Smallest card that ran it: NVIDIA T4 (16GB).
How much does it cost to run Kokoro TTS 82M in the cloud?
$0.003 per 1 hour of audio on a NVIDIA T4 at $0.14/hr, cheapest of 7 rentable cards we measured.
Can I run Kokoro TTS 82M on a 12GB, 16GB or 24GB card?
It used 1.4GB at the precision we tested. 12GB: yes; 16GB: yes; 24GB: yes.
Is the H100 80GB HBM3 or the A100 80GB SXM4 faster for Kokoro TTS 82M?
The H100 80GB HBM3: 222.0 vs 110.2 x realtime, 102% faster on our bench.

How we test

Each model runs a fixed workload on every card: a warmup, then timed runs with power, temperature and VRAM sampled every half second through NVML. Models run at the precision and settings from their model card. Datacenter cards run on Modal; consumer cards on rented machines. Non-commercially licensed models are not part of this page.