Text-to-speech · 1 model · 8 GPUs measured first-party · Updated October 2026
Which graphics card to use for text-to-speech, from first-party measurements of Kokoro TTS 82M on 8 GPUs.

244.0 x realtime on Kokoro TTS 82M, the ceiling. Measured on our bench. 48GB of VRAM, $7,500 at launch.
Text-to-speech turns written text into a spoken voice: narration for videos, audiobooks, voice agents. Kokoro is a small, high-quality open model, which makes it a good measure of how well a card handles lightweight, latency-sensitive audio work.
We measured 1 model for text-to-speech on 8 GPUs. Speed is times faster than realtime. Every number below is a first-party run on our own harness; cards absent from a model's chart have not been run on it yet.
Kokoro TTS 82M: x realtime by GPU
Which models fit which card, for text-to-speech
| Model | VRAM used | 8GB card | 12GB card | 16GB card | 24GB card | 32GB card | Licence |
|---|---|---|---|---|---|---|---|
| Kokoro TTS 82M | 1.5GB | Yes | Yes | Yes | Yes | Yes | Apache-2.0 |
From the lowest VRAM peak we measured for each model, plus 5% headroom. 'No' means it did not fit in that much memory at our settings, not that no setting ever could.
Kokoro TTS 82M: every GPU we measured
| GPU | x realtime | VRAM | Power |
|---|---|---|---|
| NVIDIA L40S | 244 | 48GB | 83.2 W |
| NVIDIA H100 80GB HBM3 | 222 | 80GB | 125.5 W |
| NVIDIA H200 | 170.8 | 141GB | 120.3 W |
| NVIDIA A100 40GB SXM4 | 142 | 40GB | 60.1 W |
| NVIDIA A100 80GB SXM4 | 110.2 | 80GB | 72.3 W |
| NVIDIA A10G | 101.1 | 24GB | 65.5 W |
| NVIDIA L4 | 97.35 | 24GB | 33.6 W |
| NVIDIA T4 | 42.88 | 16GB | 51.0 W |
What the numbers show.
Kokoro TTS 82M: fastest on the NVIDIA L40S at 244.0 x realtime, 5.69x the slowest card we measured (NVIDIA T4); it used about 1.5GB of VRAM.
How it compares. L40S: Kokoro TTS 82M 244.0 x realtime, Whisper large-v3 193.2, ACE-Step 1.5 20.44, ACE-Step v1 3.5B 13.78, DiffRhythm 2 5.43. Kokoro TTS 82M beats all 4 here.
Cost on a rented GPU. 1 hour of audio of Kokoro TTS 82M: $0.003 on a T4 ($0.14/hr, 1 min), $0.003 on a L40S ($0.79/hr, 0 min, 1.0x the cost).
Kokoro TTS 82M: cost per 1 hour of audio on rented GPUs
| GPU | Cheapest rate | Speed (x realtime) | Cost per 1 hour of audio |
|---|---|---|---|
| NVIDIA T4 | $0.14/hr | 42.88 | $0.003 |
| NVIDIA L40S | $0.79/hr | 244 | $0.003 |
| NVIDIA A100 40GB SXM4 | $0.47/hr | 142 | $0.003 |
| NVIDIA L4 | $0.44/hr | 97.35 | $0.005 |
| NVIDIA A100 80GB SXM4 | $0.95/hr | 110.2 | $0.009 |
| NVIDIA H100 80GB HBM3 | $2.14/hr | 222 | $0.010 |
| NVIDIA H200 | $3.59/hr | 170.8 | $0.021 |
Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.
Speed tiers for Kokoro TTS 82M. 100+ x realtime: 6 (L40S, H100 80GB HBM3, H200); 10-100 x realtime: 2 (L4, T4). 1x is the speed of playback.
VRAM for Kokoro TTS 82M. Measured peak 1.4GB, so 8GB is the smallest common card size; smallest card it ran on: T4 (16GB).
Power on Kokoro TTS 82M. Most efficient: L40S, 83W, 0.3 Wh per 1 hour of audio.
For text-to-speech, the NVIDIA L40S is the fastest card we measured. We have not measured a consumer card on this job yet; the picks above are datacenter and workstation hardware. Check the fit table before buying: VRAM, not speed, is what rules a card out.
Each model runs a fixed workload on every card: a warmup, then timed runs with power, temperature and VRAM sampled every half second through NVML. Models run at the precision and settings from their model card. Datacenter cards run on Modal; consumer cards on rented machines. Non-commercially licensed models are not part of this page.