NVIDIA L40, AI & Machine Learning Benchmarks & Specs

48GB · AI Score 19.5/100 · first-party measured on 12 AI workloads

19.5 AI Score Includes estimates

Every number on this page is first-party: NVIDIA L40 was run on our pinned 12-workload AI suite on 2026-07-11, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) NVIDIA L40 delivers about 135.12 tokens/sec. Stepping up to Qwen3 32B it holds roughly 34.08 tok/s. Llama 3.3 70B does not fit. It needs roughly 42GB and this card has 48GB. For image generation, SDXL runs at 5.6 it/s, and FLUX.1-dev at 1.15 it/s. 1 of the 12 workloads won't fit on 48GB at the tested precision, Llama 3.3 70B. We publish those as hard gates rather than quietly dropping to a smaller quant. NVIDIA L40 isn't a retail purchase for most people. It's rented by the hour. You can run this exact card on RunPod.

AI & Machine Learning benchmark results

Text Generation tok/s 5

Qwen3 4B211.33
Llama 3.1 8B135.12
Qwen2.5-Coder 14B74.27
Qwen3 32B34.08
WorkloadResultTelemetryData
Qwen3 4B211.33 tok/s
2.9 GB peak138 W35°C1.53 tok/WQ4_K_M
✓ Measured
Llama 3.1 8B135.12 tok/s
4.9 GB peak159 W40°C0.85 tok/WQ4_K_M
✓ Measured
Qwen2.5-Coder 14B74.27 tok/s
8.9 GB peak169 W43°C0.44 tok/WQ4_K_M
✓ Measured
Qwen3 32B34.08 tok/s
18.9 GB peak151 W45°C0.23 tok/WQ4_K_M
✓ Measured
Llama 3.3 70Bn/a tok/sestimatedEst.

Image Generation images/min 3

Stable Diffusion XL11.2
Z-Image Turbo5.625
FLUX.1 dev2.464
WorkloadResultTelemetryData
Stable Diffusion XL11.2 images/min
14.9 GB peak296 W46°C5.4 s/img
✓ Measured
Z-Image Turbo5.63 images/min
25.8 GB peak304 W52°C10.6 s/img
✓ Measured
FLUX.1 dev2.46 images/min
36.7 GB peak299 W56°C24.4 s/img
✓ Measured

Image Editing images/min 2

WorkloadResultTelemetryData
FLUX.1 Kontext dev1.11 images/min
35.5 GB peak300 W58°C53.9 s/img
✓ Measured
Qwen-Image-Edit0.48 images/min
40.2 GB peak198 W55°C126.5 s/img
✓ Measured

Video Generation frames/s 2

WorkloadResultTelemetryData
LTX-Video (distilled)2.61 frames/s
9.5 GB peak185 W48°C37.1 s/clip
✓ Measured
Wan 2.2 5B (720p)0.37 frames/s
37.3 GB peak300 W59°C133.1 s/clip
✓ Measured
How we measured this. Every result comes from our own pinned, reproducible AI suite, 12 workloads: the Qwen3-4B to Llama-70B LLM ladder (llama.cpp, Q4_K_M), SDXL / Z-Image / FLUX-dev generation, FLUX-Kontext / Qwen-Edit editing, and LTX / Wan video, run first-party on rented hardware with under 0.5% run-to-run variance. Peak VRAM, power draw, temperature and tokens-per-watt are captured per workload. “Won’t fit” rows are real data: where a model exceeds the card’s VRAM at the tested precision we record a hard gate rather than silently dropping to a smaller quant. Measured 2026-07-11 · harness 2.0.0-standalone.

NVIDIA L40 specifications

ArchitectureAda Lovelace
CUDA cores18,176
VRAM48GB GDDR6
Memory bus384-bit
Memory bandwidth864 GB/s
Boost clock2,490 MHz
TDP300 W
ProcessTSMC 4N
InterfacePCIe 4.0 x16
Release date2022-10-13
Launch MSRP$7,000

Verdict, capable, but 48GB sets the ceiling

NVIDIA L40 scores 19.5/100, #21 of 102. It ran 11 of 12; 1 exceeded its 48GB. Every figure here is our own measurement.

Relative performance: where the NVIDIA L40 lands

100% = this card, AI & Machine Learning headline metric (AI Score). #15 of 21 desktop cards in this vertical.

GPURelative%AI Score
NVIDIA A100 80GB SXM4
170%33.1
NVIDIA A800 80GB
170%33.1
NVIDIA A100 80GB PCIe
163%31.7
NVIDIA L40S
142%27.7
NVIDIA L40
100%19.5
NVIDIA A40
91%17.7
NVIDIA A100 40GB SXM4
87%17
NVIDIA A100 40GB PCIe
86%16.7
NVIDIA A10G
33%6.4

← All AI & Machine Learning GPU rankings

What this card can build

Whole-job timings, composed from our measured per-model results on this card.

WorkflowTimeEnergyBasis
24-frame storyboard7 min23.31 Whall 2 stages measured
6-panel comic page9.1 min39.9 Whall 3 stages measured
Character sheet, 12 poses12.1 min55.81 Whall 2 stages measured
60-second AI short film13 min36.37 Whall 3 stages measured
Full codebase review13.5 min37.92 Whmeasured
10 short social clips28.2 min120.12 Whall 3 stages measured
40-product photo shoot41.1 min196.9 Whall 2 stages measured
100-photo restoration batch1 h 30 min448.23 Whmeasured

Rent or buy?

This card is $7,000 to buy. The cheapest listed rate on RunPod is $0.690/hour, but that is the floor: we budget $0.828/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 8,454 GPU-hours. Below it you are paying for idle silicon.

How you would use itGPU-hours a yearRental cost a yearTime to break even
2 hours a day, hobby730$60411.6 years
8 hours a day, working on it2,920$2,4182.9 years
24/7, always-on agent8,760$7,25311.6 months

At hobby usage this card is very unlikely to pay for itself before it is superseded. Rent it. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.

Rental price

$0.690/hr+0.0% since 2026-08-14low $0.690 · high $0.690

Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.