NVIDIA L40S, AI & Machine Learning Benchmarks & Specs

48GB · AI Score 27.7/100 · first-party measured on 12 AI workloads

27.7 AI Score ✓ Measured

Every number on this page is first-party: NVIDIA L40S was run on our pinned 12-workload AI suite on 2026-07-10, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) NVIDIA L40S delivers about 134.96 tokens/sec. Stepping up to Qwen3 32B it holds roughly 34.39 tok/s. Llama 3.3 70B does not fit. It needs roughly 42GB and this card has 48GB. For image generation, SDXL runs at 8.51 it/s, and FLUX.1-dev at 1.79 it/s. 1 of the 12 workloads won't fit on 48GB at the tested precision, Llama 3.3 70B. We publish those as hard gates rather than quietly dropping to a smaller quant. NVIDIA L40S isn't a retail purchase for most people. It's rented by the hour. You can run this exact card on RunPod.

AI & Machine Learning benchmark results

Text Generation tok/s 122

SmolLM2-135M916.56
gemma-3-270m822.52
Qwen1.5-0.5B774.78
Qwen2.5-0.5B738.06
Qwen2.5-Coder-0.5B715.81
LFM2.5-1.2B661.54
Qwen3-0.6B656.26
Qwen3 0.6B655.75
Llama 3.2 1B618.87
Qwen2.5-1.5B427.23
Qwen2.5-Coder-1.5B427.12
gemma-3-1b426.49
WorkloadResultTelemetryData
SmolLM2-135M916.56 tok/s
103 W47°CQ4_K_M
✓ Measured
gemma-3-270m822.52 tok/s
94 W42°CQ4_K_M
✓ Measured
Qwen1.5-0.5B774.78 tok/s
105 W50°CQ4_K_M
✓ Measured
Qwen2.5-0.5B738.06 tok/s
100 W42°CQ4_K_M
✓ Measured
Qwen2.5-Coder-0.5B715.81 tok/s
106 W50°CQ4_K_M
✓ Measured
LFM2.5-1.2B661.54 tok/s
101 W43°CQ4_K_M
✓ Measured
Qwen3-0.6B656.26 tok/s
101 W44°CQ4_K_M
✓ Measured
Qwen3 0.6B655.75 tok/s
103 W37°CQ4_K_M
✓ Measured
Llama 3.2 1B618.87 tok/s
117 W44°CQ4_K_M
✓ Measured
Qwen2.5-1.5B427.23 tok/s
114 W45°CQ4_K_M
✓ Measured
Qwen2.5-Coder-1.5B427.12 tok/s
125 W51°CQ4_K_M
✓ Measured
gemma-3-1b426.49 tok/s
122 W44°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 1.5B426.38 tok/s
112 W45°CQ4_K_M
✓ Measured
Qwen2-1.5B416.57 tok/s
120 W49°CQ4_K_M
✓ Measured
Qwen3 1.7B405.52 tok/s
143 W40°CQ4_K_M
✓ Measured
Qwen3-1.7B405.3 tok/s
148 W52°CQ4_K_M
✓ Measured
LFM2.5-8B-A1B394.48 tok/s
125 W48°CQ4_K_M
✓ Measured
phi-2283.26 tok/s
159 W49°CQ4_K_M
✓ Measured
gemma-2-2b-it-abliterated278.14 tok/s
165 W51°CQ4_K_M
✓ Measured
gemma-2-2b278.1 tok/s
156 W47°CQ4_K_M
✓ Measured
SmolLM3-3B272.96 tok/s
172 W53°CQ4_K_M
✓ Measured
SmolLM3 3B272.88 tok/s
152 W49°CQ4_K_M
✓ Measured
Llama-3.2-3B-Instruct-uncensored271.01 tok/s
154 W54°CQ4_K_M
✓ Measured
Llama 3.2 3B270.85 tok/s
155 W45°CQ4_K_M
✓ Measured
Qwen2.5-3B270.04 tok/s
154 W47°CQ4_K_M
✓ Measured
Qwen2.5-Coder-3B269.9 tok/s
174 W52°CQ4_K_M
✓ Measured
Hermes-3-Llama-3.2-3B265.87 tok/s
164 W55°CQ4_K_M
✓ Measured
AI21-Jamba-Reasoning-3B258.41 tok/s
169 W54°CQ4_K_M
✓ Measured
DeepSeek-Coder-V2-Lite257.64 tok/s
134 W50°CQ4_K_M
✓ Measured
Dolphin X1 Trinity Nano 6B246.45 tok/s
118 W47°CQ4_K_M
✓ Measured
gpt-oss-20b233.82 tok/s
139 W51°CQ4_K_M
✓ Measured
Phi-4-mini229.28 tok/s
172 W54°CQ4_K_M
✓ Measured
Phi-3.5-mini229.06 tok/s
185 W46°CQ4_K_M
✓ Measured
Phi-4 Mini 3.8B228.93 tok/s
156 W48°CQ4_K_M
✓ Measured
Qwen3 30B A3B (Q3_K_M)224.95 tok/s
153 W58°CQ3_K_M
✓ Measured
Qwen3-Coder 30B A3B219.04 tok/s
133 W49°CQ4_K_M
✓ Measured
Qwen3-30B-A3B214.05 tok/s
126 W49°CQ4_K_M
✓ Measured
Qwen3 30B A3B213.66 tok/s
139 W49°CQ4_K_M
✓ Measured
Qwen3-4B212.05 tok/s
172 W48°CQ4_K_M
✓ Measured
Qwen3-4B-Instruct-2507212.05 tok/s
186 W52°CQ4_K_M
✓ Measured
Qwen3-4B-Thinking-2507211.88 tok/s
184 W50°CQ4_K_M
✓ Measured
Qwen3-4B-Instruct-2507210.24 tok/s
180 W57°CQ4_K_M
✓ Measured
Gemma 3 4B198.87 tok/s
159 W45°CQ4_K_M
✓ Measured
Nemotron-3-Nano-30B-A3B190.96 tok/s
139 W56°CQ4_K_M
✓ Measured
KAT-Coder-V2.5-Dev173.06 tok/s
133 W47°CQ4_K_M
✓ Measured
GLM-4.7-Flash157.89 tok/s
138 W44°CQ4_K_M
✓ Measured
Ornith-1.0-35B157.39 tok/s
135 W47°CQ4_K_M
✓ Measured
Qwen-AgentWorld-35B-A3B156.98 tok/s
135 W48°CQ4_K_M
✓ Measured
Llama-2-7B150.35 tok/s
202 W53°CQ4_K_M
✓ Measured
GLM-4.7-Flash-REAP-23B-A3B146.51 tok/s
144 W48°CQ4_K_M
✓ Measured
Mistral-7B-Instruct-v0.1144.82 tok/s
216 W54°CQ4_K_M
✓ Measured
Mistral-7B-Instruct-v0.2144.82 tok/s
212 W53°CQ4_K_M
✓ Measured
Mistral 7B v0.3144.74 tok/s
205 W50°CQ4_K_M
✓ Measured
Mistral-7B-Instruct-v0.3144.68 tok/s
209 W54°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 7B143.97 tok/s
199 W49°CQ4_K_M
✓ Measured
Qwen2.5-7B143.94 tok/s
206 W50°CQ4_K_M
✓ Measured
Qwen2.5-Coder 7B143.93 tok/s
200 W50°CQ4_K_M
✓ Measured
Qwen2.5-Coder-7B-Instruct-abliterated142.97 tok/s
209 W56°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill Llama 8B135.55 tok/s
194 W49°CQ4_K_M
✓ Measured
Dolphin 3.0 Llama 3.1 8B135.55 tok/s
203 W51°CQ4_K_M
✓ Measured
L3-8B-Stheno-v3.2135.53 tok/s
196 W53°CQ4_K_M
✓ Measured
Llama-3.1-8B135.51 tok/s
200 W55°CQ4_K_M
✓ Measured
Dolphin X1 8B135.48 tok/s
199 W51°CQ4_K_M
✓ Measured
Meta-Llama-3.1-8B135.46 tok/s
191 W49°CQ4_K_M
✓ Measured
dolphin-2.9-llama3-8b134.94 tok/s
208 W59°CQ4_K_M
✓ Measured
DarkIdol-Llama-3.1-8B-Instruct-1.2-Uncensored134.78 tok/s
217 W64°CQ4_K_M
✓ Measured
Qwen3-8B130.28 tok/s
204 W55°CQ4_K_M
✓ Measured
Qwen3 8B130.27 tok/s
200 W43°CQ4_K_M
✓ Measured
Josiefied-Qwen3-8B-abliterated-v1130.25 tok/s
203 W51°CQ4_K_M
✓ Measured
DeepSeek-R1-0528-Qwen3-8B130.23 tok/s
199 W48°CQ4_K_M
✓ Measured
Ornith-1.0-9B114.57 tok/s
198 W49°CQ4_K_M
✓ Measured
Gemma 3 12B (Q3_K_M)93.05 tok/s
235 W63°CQ3_K_M
✓ Measured
Mistral-Nemo-Instruct-240789.72 tok/s
218 W52°CQ4_K_M
✓ Measured
Phi-4 14B (Q3_K_M)89.18 tok/s
247 W64°CQ3_K_M
✓ Measured
gemma-2-9b88.87 tok/s
210 W53°CQ4_K_M
✓ Measured
NemoMix-Unleashed-12B88.85 tok/s
214 W53°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 14B (Q3_K_M)85.85 tok/s
251 W64°CQ3_K_M
✓ Measured
Gemma 3 12B80.92 tok/s
219 W48°CQ4_K_M
✓ Measured
Gemma 4 12B80.82 tok/s
211 W50°CQ4_K_M
✓ Measured
Qwen3 14B75.12 tok/s
225 W46°CQ4_K_M
✓ Measured
Hermes-4-14B75.1 tok/s
232 W54°CQ4_K_M
✓ Measured
Qwen3-14B75.01 tok/s
223 W53°CQ4_K_M
✓ Measured
Phi-4 14B74.98 tok/s
239 W53°CQ4_K_M
✓ Measured
EVA-Qwen2.5-14B-v0.274.63 tok/s
227 W53°CQ4_K_M
✓ Measured
Qwen2.5-Coder-14B-Instruct-abliterated74.63 tok/s
228 W54°CQ4_K_M
✓ Measured
Qwen2.5-14B74.62 tok/s
220 W49°CQ4_K_M
✓ Measured
Qwen2.5-Coder-14B74.6 tok/s
222 W49°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 14B74.6 tok/s
232 W53°CQ4_K_M
✓ Measured
Uncensored74.31 tok/s
236 W59°CQ4_K_M
✓ Measured
StarCoder2 15B66.88 tok/s
245 W55°CQ4_K_M
✓ Measured
Codestral 22B (Q3_K_M)59.53 tok/s
272 W62°CQ3_K_M
✓ Measured
Mistral Small 24B (Q3_K_M)58.21 tok/s
263 W64°CQ3_K_M
✓ Measured
Codestral 22B50.06 tok/s
249 W53°CQ4_K_M
✓ Measured
Mistral Small 24B48.43 tok/s
243 W54°CQ4_K_M
✓ Measured
Devstral Small 24B48.43 tok/s
247 W53°CQ4_K_M
✓ Measured
Dolphin Mistral 24B Venice48.43 tok/s
249 W54°CQ4_K_M
✓ Measured
Dolphin-Mistral-24B-Venice-Edition48.43 tok/s
240 W55°CQ4_K_M
✓ Measured
Cydonia-24B-v4.348.43 tok/s
245 W55°CQ4_K_M
✓ Measured
Dolphin 3.0 R1 Mistral 24B48.4 tok/s
244 W53°CQ4_K_M
✓ Measured
Qwen2.5-Coder 32B (Q3_K_M)41.06 tok/s
278 W66°CQ3_K_M
✓ Measured
Gemma 3 27B38.48 tok/s
248 W55°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 32B34.6 tok/s
250 W56°CQ4_K_M
✓ Measured
Qwen2.5-Coder 32B34.59 tok/s
254 W54°CQ4_K_M
✓ Measured
QwQ 32B34.59 tok/s
253 W55°CQ4_K_M
✓ Measured
Qwen2.5-32B34.59 tok/s
249 W53°CQ4_K_M
✓ Measured
Qwen3-32B34.44 tok/s
253 W54°CQ4_K_M
✓ Measured
DeepSeek-R1-Distill-Qwen-32B-abliterated34.43 tok/s
247 W58°CQ4_K_M
✓ Measured
Olmo-3.1-32B-Think34.11 tok/s
248 W54°CQ4_K_M
✓ Measured
Dolphin 2.9.1 Yi 1.5 34B33.08 tok/s
254 W55°CQ4_K_M
✓ Measured
Llama-3.3-70B16.49 tok/s
259 W58°CQ4_K_M
✓ Measured
DeepSeek-R1-Distill-Llama-70B16.49 tok/s
265 W58°CQ4_K_M
✓ Measured
Llama-3.3-70B-Instruct-abliterated16.49 tok/s
261 W56°CQ4_K_M
✓ Measured
Meta-Llama-3.1-70B16.49 tok/s
264 W58°CQ4_K_M
✓ Measured
Hermes-4-70B16.45 tok/s
276 W66°CQ4_K_M
✓ Measured
Qwen3-Coder-Next-abliterated✕ Won't fit needs ~61 GBVRAM-gated at this precision✓ Measured
Laguna-XS-2.1✕ Won't fit needs ~26 GBVRAM-gated at this precision✓ Measured
Nanbeige4.2-3B✕ Won't fit needs ~4 GBVRAM-gated at this precision✓ Measured
Qwen3-Coder-Next✕ Won't fit needs ~61 GBVRAM-gated at this precision✓ Measured
Qwen3-Next-80B-A3B-Thinking✕ Won't fit needs ~61 GBVRAM-gated at this precision✓ Measured
Qwen3-Coder-Next✕ Won't fit needs ~61 GBVRAM-gated at this precision✓ Measured
Qwen3-Next-80B-A3B-Thinking✕ Won't fit needs ~61 GBVRAM-gated at this precision✓ Measured
Qwen3-Next-80B-A3B✕ Won't fit needs ~61 GBVRAM-gated at this precision✓ Measured

Image Generation images/min 10

Sana 1.6B50.81
PixArt-Sigma XL26.95
FLUX.1 Schnell26.24
Z-Image Turbo (1024px)18.2
Stable Diffusion XL17.02
Stable Diffusion 3.5 Medium11.58
Z-Image Turbo8.325
Stable Diffusion 3.5 Large4.54
FLUX.1 dev3.836
AuraFlow v0.33.58
WorkloadResultTelemetryData
Sana 1.6B50.81 images/min
332 W46°C
✓ Measured
PixArt-Sigma XL26.95 images/min
338 W53°C
✓ Measured
FLUX.1 Schnell26.24 images/min
331 W54°C
✓ Measured
Z-Image Turbo (1024px)18.2 images/min
343 W63°C
✓ Measured
Stable Diffusion XL17.02 images/min
14.9 GB peak326 W56°C3.5 s/img
✓ Measured
Stable Diffusion 3.5 Medium11.58 images/min
346 W62°C
✓ Measured
Z-Image Turbo8.33 images/min
25.8 GB peak329 W61°C7.2 s/img
✓ Measured
Stable Diffusion 3.5 Large4.54 images/min
348 W68°C
✓ Measured
FLUX.1 dev3.84 images/min
36.7 GB peak338 W74°C15.7 s/img
✓ Measured
AuraFlow v0.33.58 images/min
348 W68°C
✓ Measured

Fine-Tuning train tok/s 4

TinyLlama 1.1B LoRA16408.1
Qwen2.5 1.5B LoRA12120.1
SmolLM2 1.7B LoRA12052.1
Qwen2.5 7B LoRA3736
WorkloadResultTelemetryData
TinyLlama 1.1B LoRA16408.1 train tok/s
246 W44°C
✓ Measured
Qwen2.5 1.5B LoRA12120.1 train tok/s
278 W46°C
✓ Measured
SmolLM2 1.7B LoRA12052.1 train tok/s
292 W48°C
✓ Measured
Qwen2.5 7B LoRA3736 train tok/s
331 W51°C
✓ Measured

LLM Serving serve tok/s 4

TinyLlama 1.1B served6278.4
Qwen2.5 1.5B served4686.8
SmolLM2 1.7B served3525.9
Qwen2.5 7B served1406.8
WorkloadResultTelemetryData
TinyLlama 1.1B served6278.4 serve tok/s
161 W37°C
✓ Measured
Qwen2.5 1.5B served4686.8 serve tok/s
203 W40°C
✓ Measured
SmolLM2 1.7B served3525.9 serve tok/s
198 W40°C
✓ Measured
Qwen2.5 7B served1406.8 serve tok/s
251 W45°C
✓ Measured

Image to Video clips/min 3

Wan 2.2 TI2V-5B (image to video)1.606
Stable Video Diffusion XT1.166
CogVideoX-5B I2V0.383
WorkloadResultTelemetryData
Wan 2.2 TI2V-5B (image to video)1.61 clips/min
333 W74°C
✓ Measured
Stable Video Diffusion XT1.17 clips/min
330 W76°C
✓ Measured
CogVideoX-5B I2V0.38 clips/min
341 W79°C
✓ Measured

Image Editing images/min 2

WorkloadResultTelemetryData
FLUX.1 Kontext dev1.69 images/min
35.5 GB peak340 W77°C35.6 s/img
✓ Measured
Qwen-Image-Edit1.06 images/min
40.2 GB peak268 W77°C56.4 s/img
✓ Measured

Video Generation frames/s 2

WorkloadResultTelemetryData
LTX-Video (distilled)8.01 frames/s
21.6 GB peak315 W64°C12.1 s/clip
✓ Measured
Wan 2.2 5B (720p)0.49 frames/s
37 GB peak321 W83°C100.1 s/clip
✓ Measured

Depth Estimation images/min 2

WorkloadResultTelemetryData
Depth Anything V2 Small970.61 images/min
85 W40°C
✓ Measured
Depth Anything V2 Large959.33 images/min
83 W40°C
✓ Measured

Segmentation images/min 2

WorkloadResultTelemetryData
SAM ViT-Base998.79 images/min
93 W45°C
✓ Measured
SAM ViT-Huge238.08 images/min
184 W50°C
✓ Measured

Vision Language images/min 2

WorkloadResultTelemetryData
Florence-2 Base184.2 images/min
99 W45°C
✓ Measured
Florence-2 Large106.54 images/min
122 W46°C
✓ Measured

Image to 3D assets/hour 1

WorkloadResultTelemetryData
TRELLIS Image-to-3D776.1 assets/hour
268 W60°C
✓ Measured

Background Removal images/min 1

WorkloadResultTelemetryData
BiRefNet863.81 images/min
92 W40°C
✓ Measured

Upscaling images/min 1

WorkloadResultTelemetryData
Swin2SR 4x Upscaler34.14 images/min
247 W50°C
✓ Measured

Text to Speech x realtime 1

WorkloadResultTelemetryData
Kokoro TTS 82M243.99 x realtime
83 W31°C
✓ Measured

Speech to Text x realtime 1

WorkloadResultTelemetryData
Whisper large-v3193.2 x realtime
97 W34°C
✓ Measured

Music Generation x realtime 1

WorkloadResultTelemetryData
MusicGen Small2.43 x realtime
158 W39°C
✓ Measured

Speculative Decoding x vs solo 1

WorkloadResultTelemetryData
Qwen2.5 1.5B + 0.5B draft0.95 x vs solo✓ Measured
How we measured this. Every result comes from our own pinned, reproducible AI suite, 12 workloads: the Qwen3-4B to Llama-70B LLM ladder (llama.cpp, Q4_K_M), SDXL / Z-Image / FLUX-dev generation, FLUX-Kontext / Qwen-Edit editing, and LTX / Wan video, run first-party on rented hardware with under 0.5% run-to-run variance. Peak VRAM, power draw, temperature and tokens-per-watt are captured per workload. “Won’t fit” rows are real data: where a model exceeds the card’s VRAM at the tested precision we record a hard gate rather than silently dropping to a smaller quant. Measured 2026-07-10 · harness 2.0.0.

NVIDIA L40S specifications

ArchitectureAda Lovelace
CUDA cores18,176
VRAM48GB GDDR6
Memory bus384-bit
Memory bandwidth864 GB/s
Boost clock2,520 MHz
TDP350 W
ProcessTSMC 4N
InterfacePCIe 4.0 x16
Release date2023-08-08
Launch MSRP$7,500

Verdict, capable, but 48GB sets the ceiling

NVIDIA L40S scores 27.7/100, #17 of 102. It ran 11 of 12; 1 exceeded its 48GB. Every figure here is our own measurement.

Relative performance: where the NVIDIA L40S lands

100% = this card, AI & Machine Learning headline metric (AI Score). #14 of 21 desktop cards in this vertical.

GPURelative%AI Score
NVIDIA H100 PCIe
168%46.4
NVIDIA A100 80GB SXM4
119%33.1
NVIDIA A800 80GB
119%33.1
NVIDIA A100 80GB PCIe
114%31.7
NVIDIA L40S
100%27.7
NVIDIA L40
70%19.5
NVIDIA A40
64%17.7
NVIDIA A100 40GB SXM4
61%17
NVIDIA A100 40GB PCIe
60%16.7

← All AI & Machine Learning GPU rankings

What this card can build

Whole-job timings, composed from our measured per-model results on this card.

WorkflowTimeEnergyBasis
50-image depth pass6 s0.07 Whmeasured
30-minute podcast pass75 s1.89 Whall 3 stages measured
500-image masking run2.2 min6.43 Whmeasured
20-asset 3D game kit2.9 min13.29 Whall 2 stages measured
24-frame storyboard4.1 min18.66 Whall 2 stages measured
60-second AI short film4.9 min22.52 Whall 3 stages measured
60-second AI short film, narrated5.1 min22.53 Whall 4 stages measured
200-product catalogue cutout6.2 min24.48 Whall 2 stages measured
6-panel comic page6.5 min30.3 Whall 3 stages measured
Character sheet, 12 poses8.4 min41.61 Whall 2 stages measured
Full codebase review13.5 min49.6 Whmeasured
10 short social clips19.1 min96.77 Whall 3 stages measured
40-product photo shoot26.7 min146.58 Whall 2 stages measured
40-product shoot, start to finish28 min151.48 Whall 4 stages measured
20 long-form articles28.5 min122.26 Whmeasured
100-photo restoration batch59.6 min334.51 Whmeasured
100-photo restore and enlarge62.5 min346.58 Whall 2 stages measured

Rent or buy?

This card is $7,500 to buy. The cheapest listed rate on RunPod is $0.790/hour, but that is the floor: we budget $0.948/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 7,911 GPU-hours. Below it you are paying for idle silicon.

How you would use itGPU-hours a yearRental cost a yearTime to break even
2 hours a day, hobby730$69210.8 years
8 hours a day, working on it2,920$2,7682.7 years
24/7, always-on agent8,760$8,30410.8 months

At hobby usage this card is very unlikely to pay for itself before it is superseded. Rent it. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.

Rental price

$0.790/hr+0.0% since 2026-08-14low $0.790 · high $0.790

Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.