Speech-to-text with Whisper · 1 model · 8 GPUs measured first-party · Updated October 2026

Best GPU for Speech-to-text with Whisper

Which graphics card to use for speech-to-text with Whisper, from first-party measurements of Whisper large-v3 on 8 GPUs.

Fastest we measured
NVIDIA L40S

NVIDIA L40S

193.2 x realtime on Whisper large-v3, the ceiling. Measured on our bench. 48GB of VRAM, $7,500 at launch.

Pros
  • 193.2 x realtime on Whisper large-v3
  • 48GB, clears the Whisper large-v3 floor
  • Rentable by the hour rather than bought
Cons
  • 350W board rating
  • Datacenter or workstation hardware, not a retail purchase
1
Models measured
Whisper large-v3
8
GPUs measured
first-party runs, not spec-sheet estimates
193.2x realtime
Fastest: NVIDIA L40S
on Whisper large-v3
5.3GB
Lightest model's VRAM need
measured peak, +5% headroom

Whisper turns audio into text: podcasts, meetings, video subtitles. It is one of the most-used open models there is, and because it runs much faster than realtime on most cards the question is less 'can it run' and more 'how many hours of audio per hour of GPU'.

We measured 1 model for speech-to-text with Whisper on 8 GPUs. Speed is times faster than realtime (60x = one hour of audio in one minute). Every number below is a first-party run on our own harness; cards absent from a model's chart have not been run on it yet.

Whisper large-v3: x realtime by GPU

NVIDIA L40S
193.2 x realtime
NVIDIA H100 80GB HBM3
181.5 x realtime
NVIDIA H200
163.8 x realtime
NVIDIA A100 40GB SXM4
102.5 x realtime
NVIDIA A10G
86.07 x realtime
NVIDIA A100 80GB SXM4
73.83 x realtime
NVIDIA L4
70.11 x realtime
NVIDIA T4
44.36 x realtime

Which models fit which card, for speech-to-text with Whisper

ModelVRAM used8GB card12GB card16GB card24GB card32GB cardLicence
Whisper large-v35.3GBYesYesYesYesYesMIT

From the lowest VRAM peak we measured for each model, plus 5% headroom. 'No' means it did not fit in that much memory at our settings, not that no setting ever could.

Whisper large-v3: every GPU we measured

NVIDIA L40S193.2
NVIDIA H100 80GB HBM3181.5
NVIDIA H200163.8
NVIDIA A100 40GB SXM4102.5
NVIDIA A10G86.07
NVIDIA A100 80GB SXM473.83
NVIDIA L470.11
NVIDIA T444.36
GPUx realtimeVRAMPower
NVIDIA L40S193.248GB96.6 W
NVIDIA H100 80GB HBM3181.580GB137.6 W
NVIDIA H200163.8141GB159.0 W
NVIDIA A100 40GB SXM4102.540GB189.5 W
NVIDIA A10G86.0724GB116.0 W
NVIDIA A100 80GB SXM473.8380GB151.1 W
NVIDIA L470.1124GB55.7 W
NVIDIA T444.3616GB64.6 W

What the numbers show.

Whisper large-v3: fastest on the NVIDIA L40S at 193.2 x realtime, 4.36x the slowest card we measured (NVIDIA T4); it used about 5.3GB of VRAM.

How it compares. L40S: Whisper large-v3 193.2 x realtime, Kokoro TTS 82M 244.0, ACE-Step 1.5 20.44, ACE-Step v1 3.5B 13.78, DiffRhythm 2 5.43. 1 of 4 beat Whisper large-v3 here.

Cost on a rented GPU. 1 hour of audio of Whisper large-v3: $0.003 on a T4 ($0.14/hr, 1 min), $0.004 on a L40S ($0.79/hr, 0 min, 1.3x the cost).

Whisper large-v3: cost per 1 hour of audio on rented GPUs

NVIDIA T4$0.14/hr
NVIDIA L40S$0.79/hr
NVIDIA A100 40GB SXM4$0.47/hr
NVIDIA L4$0.44/hr
NVIDIA H100 80GB HBM3$2.14/hr
NVIDIA A100 80GB SXM4$0.95/hr
NVIDIA H200$3.59/hr
GPUCheapest rateSpeed (x realtime)Cost per 1 hour of audio
NVIDIA T4$0.14/hr44.36$0.003
NVIDIA L40S$0.79/hr193.2$0.004
NVIDIA A100 40GB SXM4$0.47/hr102.5$0.005
NVIDIA L4$0.44/hr70.11$0.006
NVIDIA H100 80GB HBM3$2.14/hr181.5$0.012
NVIDIA A100 80GB SXM4$0.95/hr73.83$0.013
NVIDIA H200$3.59/hr163.8$0.022

Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.

Speed tiers for Whisper large-v3. 100+ x realtime: 4 (L40S, H100 80GB HBM3, H200); 10-100 x realtime: 4 (A10G, A100 80GB SXM4, L4). 1x is the speed of playback.

VRAM for Whisper large-v3. Measured peak 5.1GB, so 8GB is the smallest common card size; smallest card it ran on: T4 (16GB).

Power on Whisper large-v3. Most efficient: L40S, 97W, 0.5 Wh per 1 hour of audio. Hungriest: A100 40GB SXM4, 190W, 1.8 Wh.

Our verdict

For speech-to-text with Whisper, the NVIDIA L40S is the fastest card we measured. We have not measured a consumer card on this job yet; the picks above are datacenter and workstation hardware. Check the fit table before buying: VRAM, not speed, is what rules a card out.

FAQ

What is the fastest GPU for speech-to-text with Whisper?
In our runs, the NVIDIA L40S at 193.2 x realtime on Whisper large-v3. We measured 1 model on 8 GPUs for this page.
How much VRAM do I need for speech-to-text with Whisper?
The lightest model here, Whisper large-v3, used about 5.3GB. The table above shows which models fit 8, 12, 16, 24 and 32GB cards, from measured peaks.
Are these numbers measured or estimated?
Measured. Every number on this page is a first-party run on our own harness, with power and VRAM sampled during the run. Cards we have not run yet are simply absent, not filled in.
What GPU do I need to run Whisper large-v3?
About 5GB. Smallest card that ran it: NVIDIA T4 (16GB).
How much does it cost to run Whisper large-v3 in the cloud?
$0.003 per 1 hour of audio on a NVIDIA T4 at $0.14/hr, cheapest of 7 rentable cards we measured.
Can I run Whisper large-v3 on a 12GB, 16GB or 24GB card?
It used 5.1GB at the precision we tested. 12GB: yes; 16GB: yes; 24GB: yes.
Is the H100 80GB HBM3 or the A100 80GB SXM4 faster for Whisper large-v3?
The H100 80GB HBM3: 181.5 vs 73.83 x realtime, 146% faster on our bench.

How we test

Each model runs a fixed workload on every card: a warmup, then timed runs with power, temperature and VRAM sampled every half second through NVML. Models run at the precision and settings from their model card. Datacenter cards run on Modal; consumer cards on rented machines. Non-commercially licensed models are not part of this page.