Speech-to-text with Whisper · 1 model · 8 GPUs measured first-party · Updated October 2026
Which graphics card to use for speech-to-text with Whisper, from first-party measurements of Whisper large-v3 on 8 GPUs.

193.2 x realtime on Whisper large-v3, the ceiling. Measured on our bench. 48GB of VRAM, $7,500 at launch.
Whisper turns audio into text: podcasts, meetings, video subtitles. It is one of the most-used open models there is, and because it runs much faster than realtime on most cards the question is less 'can it run' and more 'how many hours of audio per hour of GPU'.
We measured 1 model for speech-to-text with Whisper on 8 GPUs. Speed is times faster than realtime (60x = one hour of audio in one minute). Every number below is a first-party run on our own harness; cards absent from a model's chart have not been run on it yet.
Whisper large-v3: x realtime by GPU
Which models fit which card, for speech-to-text with Whisper
| Model | VRAM used | 8GB card | 12GB card | 16GB card | 24GB card | 32GB card | Licence |
|---|---|---|---|---|---|---|---|
| Whisper large-v3 | 5.3GB | Yes | Yes | Yes | Yes | Yes | MIT |
From the lowest VRAM peak we measured for each model, plus 5% headroom. 'No' means it did not fit in that much memory at our settings, not that no setting ever could.
Whisper large-v3: every GPU we measured
| GPU | x realtime | VRAM | Power |
|---|---|---|---|
| NVIDIA L40S | 193.2 | 48GB | 96.6 W |
| NVIDIA H100 80GB HBM3 | 181.5 | 80GB | 137.6 W |
| NVIDIA H200 | 163.8 | 141GB | 159.0 W |
| NVIDIA A100 40GB SXM4 | 102.5 | 40GB | 189.5 W |
| NVIDIA A10G | 86.07 | 24GB | 116.0 W |
| NVIDIA A100 80GB SXM4 | 73.83 | 80GB | 151.1 W |
| NVIDIA L4 | 70.11 | 24GB | 55.7 W |
| NVIDIA T4 | 44.36 | 16GB | 64.6 W |
What the numbers show.
Whisper large-v3: fastest on the NVIDIA L40S at 193.2 x realtime, 4.36x the slowest card we measured (NVIDIA T4); it used about 5.3GB of VRAM.
How it compares. L40S: Whisper large-v3 193.2 x realtime, Kokoro TTS 82M 244.0, ACE-Step 1.5 20.44, ACE-Step v1 3.5B 13.78, DiffRhythm 2 5.43. 1 of 4 beat Whisper large-v3 here.
Cost on a rented GPU. 1 hour of audio of Whisper large-v3: $0.003 on a T4 ($0.14/hr, 1 min), $0.004 on a L40S ($0.79/hr, 0 min, 1.3x the cost).
Whisper large-v3: cost per 1 hour of audio on rented GPUs
| GPU | Cheapest rate | Speed (x realtime) | Cost per 1 hour of audio |
|---|---|---|---|
| NVIDIA T4 | $0.14/hr | 44.36 | $0.003 |
| NVIDIA L40S | $0.79/hr | 193.2 | $0.004 |
| NVIDIA A100 40GB SXM4 | $0.47/hr | 102.5 | $0.005 |
| NVIDIA L4 | $0.44/hr | 70.11 | $0.006 |
| NVIDIA H100 80GB HBM3 | $2.14/hr | 181.5 | $0.012 |
| NVIDIA A100 80GB SXM4 | $0.95/hr | 73.83 | $0.013 |
| NVIDIA H200 | $3.59/hr | 163.8 | $0.022 |
Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.
Speed tiers for Whisper large-v3. 100+ x realtime: 4 (L40S, H100 80GB HBM3, H200); 10-100 x realtime: 4 (A10G, A100 80GB SXM4, L4). 1x is the speed of playback.
VRAM for Whisper large-v3. Measured peak 5.1GB, so 8GB is the smallest common card size; smallest card it ran on: T4 (16GB).
Power on Whisper large-v3. Most efficient: L40S, 97W, 0.5 Wh per 1 hour of audio. Hungriest: A100 40GB SXM4, 190W, 1.8 Wh.
For speech-to-text with Whisper, the NVIDIA L40S is the fastest card we measured. We have not measured a consumer card on this job yet; the picks above are datacenter and workstation hardware. Check the fit table before buying: VRAM, not speed, is what rules a card out.
Each model runs a fixed workload on every card: a warmup, then timed runs with power, temperature and VRAM sampled every half second through NVML. Models run at the precision and settings from their model card. Datacenter cards run on Modal; consumer cards on rented machines. Non-commercially licensed models are not part of this page.