48GB · AI Score 27.7/100 · first-party measured on 12 AI workloads
27.7 AI Score ✓ Measured
Every number on this page is first-party: NVIDIA L40S was run on our pinned 12-workload AI suite on 2026-07-10, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) NVIDIA L40S delivers about 134.96 tokens/sec. Stepping up to Qwen3 32B it holds roughly 34.39 tok/s. Llama 3.3 70B does not fit. It needs roughly 42GB and this card has 48GB. For image generation, SDXL runs at 8.51 it/s, and FLUX.1-dev at 1.79 it/s. 1 of the 12 workloads won't fit on 48GB at the tested precision, Llama 3.3 70B. We publish those as hard gates rather than quietly dropping to a smaller quant. NVIDIA L40S isn't a retail purchase for most people. It's rented by the hour. You can run this exact card on RunPod.
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| SmolLM2-135M | 916.56 tok/s | 103 W47°CQ4_K_M | ✓ Measured |
| gemma-3-270m | 822.52 tok/s | 94 W42°CQ4_K_M | ✓ Measured |
| Qwen1.5-0.5B | 774.78 tok/s | 105 W50°CQ4_K_M | ✓ Measured |
| Qwen2.5-0.5B | 738.06 tok/s | 100 W42°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder-0.5B | 715.81 tok/s | 106 W50°CQ4_K_M | ✓ Measured |
| LFM2.5-1.2B | 661.54 tok/s | 101 W43°CQ4_K_M | ✓ Measured |
| Qwen3-0.6B | 656.26 tok/s | 101 W44°CQ4_K_M | ✓ Measured |
| Qwen3 0.6B | 655.75 tok/s | 103 W37°CQ4_K_M | ✓ Measured |
| Llama 3.2 1B | 618.87 tok/s | 117 W44°CQ4_K_M | ✓ Measured |
| Qwen2.5-1.5B | 427.23 tok/s | 114 W45°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder-1.5B | 427.12 tok/s | 125 W51°CQ4_K_M | ✓ Measured |
| gemma-3-1b | 426.49 tok/s | 122 W44°CQ4_K_M | ✓ Measured |
| DeepSeek-R1 Distill 1.5B | 426.38 tok/s | 112 W45°CQ4_K_M | ✓ Measured |
| Qwen2-1.5B | 416.57 tok/s | 120 W49°CQ4_K_M | ✓ Measured |
| Qwen3 1.7B | 405.52 tok/s | 143 W40°CQ4_K_M | ✓ Measured |
| Qwen3-1.7B | 405.3 tok/s | 148 W52°CQ4_K_M | ✓ Measured |
| LFM2.5-8B-A1B | 394.48 tok/s | 125 W48°CQ4_K_M | ✓ Measured |
| phi-2 | 283.26 tok/s | 159 W49°CQ4_K_M | ✓ Measured |
| gemma-2-2b-it-abliterated | 278.14 tok/s | 165 W51°CQ4_K_M | ✓ Measured |
| gemma-2-2b | 278.1 tok/s | 156 W47°CQ4_K_M | ✓ Measured |
| SmolLM3-3B | 272.96 tok/s | 172 W53°CQ4_K_M | ✓ Measured |
| SmolLM3 3B | 272.88 tok/s | 152 W49°CQ4_K_M | ✓ Measured |
| Llama-3.2-3B-Instruct-uncensored | 271.01 tok/s | 154 W54°CQ4_K_M | ✓ Measured |
| Llama 3.2 3B | 270.85 tok/s | 155 W45°CQ4_K_M | ✓ Measured |
| Qwen2.5-3B | 270.04 tok/s | 154 W47°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder-3B | 269.9 tok/s | 174 W52°CQ4_K_M | ✓ Measured |
| Hermes-3-Llama-3.2-3B | 265.87 tok/s | 164 W55°CQ4_K_M | ✓ Measured |
| AI21-Jamba-Reasoning-3B | 258.41 tok/s | 169 W54°CQ4_K_M | ✓ Measured |
| DeepSeek-Coder-V2-Lite | 257.64 tok/s | 134 W50°CQ4_K_M | ✓ Measured |
| Dolphin X1 Trinity Nano 6B | 246.45 tok/s | 118 W47°CQ4_K_M | ✓ Measured |
| gpt-oss-20b | 233.82 tok/s | 139 W51°CQ4_K_M | ✓ Measured |
| Phi-4-mini | 229.28 tok/s | 172 W54°CQ4_K_M | ✓ Measured |
| Phi-3.5-mini | 229.06 tok/s | 185 W46°CQ4_K_M | ✓ Measured |
| Phi-4 Mini 3.8B | 228.93 tok/s | 156 W48°CQ4_K_M | ✓ Measured |
| Qwen3 30B A3B (Q3_K_M) | 224.95 tok/s | 153 W58°CQ3_K_M | ✓ Measured |
| Qwen3-Coder 30B A3B | 219.04 tok/s | 133 W49°CQ4_K_M | ✓ Measured |
| Qwen3-30B-A3B | 214.05 tok/s | 126 W49°CQ4_K_M | ✓ Measured |
| Qwen3 30B A3B | 213.66 tok/s | 139 W49°CQ4_K_M | ✓ Measured |
| Qwen3-4B | 212.05 tok/s | 172 W48°CQ4_K_M | ✓ Measured |
| Qwen3-4B-Instruct-2507 | 212.05 tok/s | 186 W52°CQ4_K_M | ✓ Measured |
| Qwen3-4B-Thinking-2507 | 211.88 tok/s | 184 W50°CQ4_K_M | ✓ Measured |
| Qwen3-4B-Instruct-2507 | 210.24 tok/s | 180 W57°CQ4_K_M | ✓ Measured |
| Gemma 3 4B | 198.87 tok/s | 159 W45°CQ4_K_M | ✓ Measured |
| Nemotron-3-Nano-30B-A3B | 190.96 tok/s | 139 W56°CQ4_K_M | ✓ Measured |
| KAT-Coder-V2.5-Dev | 173.06 tok/s | 133 W47°CQ4_K_M | ✓ Measured |
| GLM-4.7-Flash | 157.89 tok/s | 138 W44°CQ4_K_M | ✓ Measured |
| Ornith-1.0-35B | 157.39 tok/s | 135 W47°CQ4_K_M | ✓ Measured |
| Qwen-AgentWorld-35B-A3B | 156.98 tok/s | 135 W48°CQ4_K_M | ✓ Measured |
| Llama-2-7B | 150.35 tok/s | 202 W53°CQ4_K_M | ✓ Measured |
| GLM-4.7-Flash-REAP-23B-A3B | 146.51 tok/s | 144 W48°CQ4_K_M | ✓ Measured |
| Mistral-7B-Instruct-v0.1 | 144.82 tok/s | 216 W54°CQ4_K_M | ✓ Measured |
| Mistral-7B-Instruct-v0.2 | 144.82 tok/s | 212 W53°CQ4_K_M | ✓ Measured |
| Mistral 7B v0.3 | 144.74 tok/s | 205 W50°CQ4_K_M | ✓ Measured |
| Mistral-7B-Instruct-v0.3 | 144.68 tok/s | 209 W54°CQ4_K_M | ✓ Measured |
| DeepSeek-R1 Distill 7B | 143.97 tok/s | 199 W49°CQ4_K_M | ✓ Measured |
| Qwen2.5-7B | 143.94 tok/s | 206 W50°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder 7B | 143.93 tok/s | 200 W50°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder-7B-Instruct-abliterated | 142.97 tok/s | 209 W56°CQ4_K_M | ✓ Measured |
| DeepSeek-R1 Distill Llama 8B | 135.55 tok/s | 194 W49°CQ4_K_M | ✓ Measured |
| Dolphin 3.0 Llama 3.1 8B | 135.55 tok/s | 203 W51°CQ4_K_M | ✓ Measured |
| L3-8B-Stheno-v3.2 | 135.53 tok/s | 196 W53°CQ4_K_M | ✓ Measured |
| Llama-3.1-8B | 135.51 tok/s | 200 W55°CQ4_K_M | ✓ Measured |
| Dolphin X1 8B | 135.48 tok/s | 199 W51°CQ4_K_M | ✓ Measured |
| Meta-Llama-3.1-8B | 135.46 tok/s | 191 W49°CQ4_K_M | ✓ Measured |
| dolphin-2.9-llama3-8b | 134.94 tok/s | 208 W59°CQ4_K_M | ✓ Measured |
| DarkIdol-Llama-3.1-8B-Instruct-1.2-Uncensored | 134.78 tok/s | 217 W64°CQ4_K_M | ✓ Measured |
| Qwen3-8B | 130.28 tok/s | 204 W55°CQ4_K_M | ✓ Measured |
| Qwen3 8B | 130.27 tok/s | 200 W43°CQ4_K_M | ✓ Measured |
| Josiefied-Qwen3-8B-abliterated-v1 | 130.25 tok/s | 203 W51°CQ4_K_M | ✓ Measured |
| DeepSeek-R1-0528-Qwen3-8B | 130.23 tok/s | 199 W48°CQ4_K_M | ✓ Measured |
| Ornith-1.0-9B | 114.57 tok/s | 198 W49°CQ4_K_M | ✓ Measured |
| Gemma 3 12B (Q3_K_M) | 93.05 tok/s | 235 W63°CQ3_K_M | ✓ Measured |
| Mistral-Nemo-Instruct-2407 | 89.72 tok/s | 218 W52°CQ4_K_M | ✓ Measured |
| Phi-4 14B (Q3_K_M) | 89.18 tok/s | 247 W64°CQ3_K_M | ✓ Measured |
| gemma-2-9b | 88.87 tok/s | 210 W53°CQ4_K_M | ✓ Measured |
| NemoMix-Unleashed-12B | 88.85 tok/s | 214 W53°CQ4_K_M | ✓ Measured |
| DeepSeek-R1 Distill 14B (Q3_K_M) | 85.85 tok/s | 251 W64°CQ3_K_M | ✓ Measured |
| Gemma 3 12B | 80.92 tok/s | 219 W48°CQ4_K_M | ✓ Measured |
| Gemma 4 12B | 80.82 tok/s | 211 W50°CQ4_K_M | ✓ Measured |
| Qwen3 14B | 75.12 tok/s | 225 W46°CQ4_K_M | ✓ Measured |
| Hermes-4-14B | 75.1 tok/s | 232 W54°CQ4_K_M | ✓ Measured |
| Qwen3-14B | 75.01 tok/s | 223 W53°CQ4_K_M | ✓ Measured |
| Phi-4 14B | 74.98 tok/s | 239 W53°CQ4_K_M | ✓ Measured |
| EVA-Qwen2.5-14B-v0.2 | 74.63 tok/s | 227 W53°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder-14B-Instruct-abliterated | 74.63 tok/s | 228 W54°CQ4_K_M | ✓ Measured |
| Qwen2.5-14B | 74.62 tok/s | 220 W49°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder-14B | 74.6 tok/s | 222 W49°CQ4_K_M | ✓ Measured |
| DeepSeek-R1 Distill 14B | 74.6 tok/s | 232 W53°CQ4_K_M | ✓ Measured |
| Uncensored | 74.31 tok/s | 236 W59°CQ4_K_M | ✓ Measured |
| StarCoder2 15B | 66.88 tok/s | 245 W55°CQ4_K_M | ✓ Measured |
| Codestral 22B (Q3_K_M) | 59.53 tok/s | 272 W62°CQ3_K_M | ✓ Measured |
| Mistral Small 24B (Q3_K_M) | 58.21 tok/s | 263 W64°CQ3_K_M | ✓ Measured |
| Codestral 22B | 50.06 tok/s | 249 W53°CQ4_K_M | ✓ Measured |
| Mistral Small 24B | 48.43 tok/s | 243 W54°CQ4_K_M | ✓ Measured |
| Devstral Small 24B | 48.43 tok/s | 247 W53°CQ4_K_M | ✓ Measured |
| Dolphin Mistral 24B Venice | 48.43 tok/s | 249 W54°CQ4_K_M | ✓ Measured |
| Dolphin-Mistral-24B-Venice-Edition | 48.43 tok/s | 240 W55°CQ4_K_M | ✓ Measured |
| Cydonia-24B-v4.3 | 48.43 tok/s | 245 W55°CQ4_K_M | ✓ Measured |
| Dolphin 3.0 R1 Mistral 24B | 48.4 tok/s | 244 W53°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder 32B (Q3_K_M) | 41.06 tok/s | 278 W66°CQ3_K_M | ✓ Measured |
| Gemma 3 27B | 38.48 tok/s | 248 W55°CQ4_K_M | ✓ Measured |
| DeepSeek-R1 Distill 32B | 34.6 tok/s | 250 W56°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder 32B | 34.59 tok/s | 254 W54°CQ4_K_M | ✓ Measured |
| QwQ 32B | 34.59 tok/s | 253 W55°CQ4_K_M | ✓ Measured |
| Qwen2.5-32B | 34.59 tok/s | 249 W53°CQ4_K_M | ✓ Measured |
| Qwen3-32B | 34.44 tok/s | 253 W54°CQ4_K_M | ✓ Measured |
| DeepSeek-R1-Distill-Qwen-32B-abliterated | 34.43 tok/s | 247 W58°CQ4_K_M | ✓ Measured |
| Olmo-3.1-32B-Think | 34.11 tok/s | 248 W54°CQ4_K_M | ✓ Measured |
| Dolphin 2.9.1 Yi 1.5 34B | 33.08 tok/s | 254 W55°CQ4_K_M | ✓ Measured |
| Llama-3.3-70B | 16.49 tok/s | 259 W58°CQ4_K_M | ✓ Measured |
| DeepSeek-R1-Distill-Llama-70B | 16.49 tok/s | 265 W58°CQ4_K_M | ✓ Measured |
| Llama-3.3-70B-Instruct-abliterated | 16.49 tok/s | 261 W56°CQ4_K_M | ✓ Measured |
| Meta-Llama-3.1-70B | 16.49 tok/s | 264 W58°CQ4_K_M | ✓ Measured |
| Hermes-4-70B | 16.45 tok/s | 276 W66°CQ4_K_M | ✓ Measured |
| Qwen3-Coder-Next-abliterated | ✕ Won't fit needs ~61 GB | VRAM-gated at this precision | ✓ Measured |
| Laguna-XS-2.1 | ✕ Won't fit needs ~26 GB | VRAM-gated at this precision | ✓ Measured |
| Nanbeige4.2-3B | ✕ Won't fit needs ~4 GB | VRAM-gated at this precision | ✓ Measured |
| Qwen3-Coder-Next | ✕ Won't fit needs ~61 GB | VRAM-gated at this precision | ✓ Measured |
| Qwen3-Next-80B-A3B-Thinking | ✕ Won't fit needs ~61 GB | VRAM-gated at this precision | ✓ Measured |
| Qwen3-Coder-Next | ✕ Won't fit needs ~61 GB | VRAM-gated at this precision | ✓ Measured |
| Qwen3-Next-80B-A3B-Thinking | ✕ Won't fit needs ~61 GB | VRAM-gated at this precision | ✓ Measured |
| Qwen3-Next-80B-A3B | ✕ Won't fit needs ~61 GB | VRAM-gated at this precision | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Sana 1.6B | 50.81 images/min | 332 W46°C | ✓ Measured |
| PixArt-Sigma XL | 26.95 images/min | 338 W53°C | ✓ Measured |
| FLUX.1 Schnell | 26.24 images/min | 331 W54°C | ✓ Measured |
| Z-Image Turbo (1024px) | 18.2 images/min | 343 W63°C | ✓ Measured |
| Stable Diffusion XL | 17.02 images/min | 14.9 GB peak326 W56°C3.5 s/img | ✓ Measured |
| Stable Diffusion 3.5 Medium | 11.58 images/min | 346 W62°C | ✓ Measured |
| Z-Image Turbo | 8.33 images/min | 25.8 GB peak329 W61°C7.2 s/img | ✓ Measured |
| Stable Diffusion 3.5 Large | 4.54 images/min | 348 W68°C | ✓ Measured |
| FLUX.1 dev | 3.84 images/min | 36.7 GB peak338 W74°C15.7 s/img | ✓ Measured |
| AuraFlow v0.3 | 3.58 images/min | 348 W68°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| TinyLlama 1.1B LoRA | 16408.1 train tok/s | 246 W44°C | ✓ Measured |
| Qwen2.5 1.5B LoRA | 12120.1 train tok/s | 278 W46°C | ✓ Measured |
| SmolLM2 1.7B LoRA | 12052.1 train tok/s | 292 W48°C | ✓ Measured |
| Qwen2.5 7B LoRA | 3736 train tok/s | 331 W51°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| TinyLlama 1.1B served | 6278.4 serve tok/s | 161 W37°C | ✓ Measured |
| Qwen2.5 1.5B served | 4686.8 serve tok/s | 203 W40°C | ✓ Measured |
| SmolLM2 1.7B served | 3525.9 serve tok/s | 198 W40°C | ✓ Measured |
| Qwen2.5 7B served | 1406.8 serve tok/s | 251 W45°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Wan 2.2 TI2V-5B (image to video) | 1.61 clips/min | 333 W74°C | ✓ Measured |
| Stable Video Diffusion XT | 1.17 clips/min | 330 W76°C | ✓ Measured |
| CogVideoX-5B I2V | 0.38 clips/min | 341 W79°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| FLUX.1 Kontext dev | 1.69 images/min | 35.5 GB peak340 W77°C35.6 s/img | ✓ Measured |
| Qwen-Image-Edit | 1.06 images/min | 40.2 GB peak268 W77°C56.4 s/img | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| LTX-Video (distilled) | 8.01 frames/s | 21.6 GB peak315 W64°C12.1 s/clip | ✓ Measured |
| Wan 2.2 5B (720p) | 0.49 frames/s | 37 GB peak321 W83°C100.1 s/clip | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Depth Anything V2 Small | 970.61 images/min | 85 W40°C | ✓ Measured |
| Depth Anything V2 Large | 959.33 images/min | 83 W40°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| SAM ViT-Base | 998.79 images/min | 93 W45°C | ✓ Measured |
| SAM ViT-Huge | 238.08 images/min | 184 W50°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Florence-2 Base | 184.2 images/min | 99 W45°C | ✓ Measured |
| Florence-2 Large | 106.54 images/min | 122 W46°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| TRELLIS Image-to-3D | 776.1 assets/hour | 268 W60°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| BiRefNet | 863.81 images/min | 92 W40°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Swin2SR 4x Upscaler | 34.14 images/min | 247 W50°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Kokoro TTS 82M | 243.99 x realtime | 83 W31°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Whisper large-v3 | 193.2 x realtime | 97 W34°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| MusicGen Small | 2.43 x realtime | 158 W39°C | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Qwen2.5 1.5B + 0.5B draft | 0.95 x vs solo | ✓ Measured |
| Architecture | Ada Lovelace |
| CUDA cores | 18,176 |
| VRAM | 48GB GDDR6 |
| Memory bus | 384-bit |
| Memory bandwidth | 864 GB/s |
| Boost clock | 2,520 MHz |
| TDP | 350 W |
| Process | TSMC 4N |
| Interface | PCIe 4.0 x16 |
| Release date | 2023-08-08 |
| Launch MSRP | $7,500 |
NVIDIA L40S scores 27.7/100, #17 of 102. It ran 11 of 12; 1 exceeded its 48GB. Every figure here is our own measurement.
100% = this card, AI & Machine Learning headline metric (AI Score). #14 of 21 desktop cards in this vertical.
| GPU | Relative | % | AI Score |
|---|---|---|---|
| NVIDIA H100 PCIe | 168% | 46.4 | |
| NVIDIA A100 80GB SXM4 | 119% | 33.1 | |
| NVIDIA A800 80GB | 119% | 33.1 | |
| NVIDIA A100 80GB PCIe | 114% | 31.7 | |
| NVIDIA L40S | 100% | 27.7 | |
| NVIDIA L40 | 70% | 19.5 | |
| NVIDIA A40 | 64% | 17.7 | |
| NVIDIA A100 40GB SXM4 | 61% | 17 | |
| NVIDIA A100 40GB PCIe | 60% | 16.7 |
← All AI & Machine Learning GPU rankings
Whole-job timings, composed from our measured per-model results on this card.
| Workflow | Time | Energy | Basis |
|---|---|---|---|
| 50-image depth pass | 6 s | 0.07 Wh | measured |
| 30-minute podcast pass | 75 s | 1.89 Wh | all 3 stages measured |
| 500-image masking run | 2.2 min | 6.43 Wh | measured |
| 20-asset 3D game kit | 2.9 min | 13.29 Wh | all 2 stages measured |
| 24-frame storyboard | 4.1 min | 18.66 Wh | all 2 stages measured |
| 60-second AI short film | 4.9 min | 22.52 Wh | all 3 stages measured |
| 60-second AI short film, narrated | 5.1 min | 22.53 Wh | all 4 stages measured |
| 200-product catalogue cutout | 6.2 min | 24.48 Wh | all 2 stages measured |
| 6-panel comic page | 6.5 min | 30.3 Wh | all 3 stages measured |
| Character sheet, 12 poses | 8.4 min | 41.61 Wh | all 2 stages measured |
| Full codebase review | 13.5 min | 49.6 Wh | measured |
| 10 short social clips | 19.1 min | 96.77 Wh | all 3 stages measured |
| 40-product photo shoot | 26.7 min | 146.58 Wh | all 2 stages measured |
| 40-product shoot, start to finish | 28 min | 151.48 Wh | all 4 stages measured |
| 20 long-form articles | 28.5 min | 122.26 Wh | measured |
| 100-photo restoration batch | 59.6 min | 334.51 Wh | measured |
| 100-photo restore and enlarge | 62.5 min | 346.58 Wh | all 2 stages measured |
This card is $7,500 to buy. The cheapest listed rate on RunPod is $0.790/hour, but that is the floor: we budget $0.948/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 7,911 GPU-hours. Below it you are paying for idle silicon.
| How you would use it | GPU-hours a year | Rental cost a year | Time to break even |
|---|---|---|---|
| 2 hours a day, hobby | 730 | $692 | 10.8 years |
| 8 hours a day, working on it | 2,920 | $2,768 | 2.7 years |
| 24/7, always-on agent | 8,760 | $8,304 | 10.8 months |
At hobby usage this card is very unlikely to pay for itself before it is superseded. Rent it. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.
Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.