NVIDIA L4, AI & Machine Learning Benchmarks & Specs

24GB · AI Score 5.0/100 · first-party measured on 12 AI workloads

5 AI Score ✓ Measured

Every number on this page is first-party: NVIDIA L4 was run on our pinned 12-workload AI suite on 2026-07-10, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) NVIDIA L4 delivers about 50.45 tokens/sec. Llama 3.3 70B does not fit. It needs roughly 42GB and this card has 24GB. For image generation, SDXL runs at 2.59 it/s, while FLUX.1-dev won't fit at BF16 (needs ~26GB). 5 of the 12 workloads won't fit on 24GB at the tested precision, Qwen3 32B, Llama 3.3 70B, FLUX.1-dev, FLUX.1 Kontext and others. We publish those as hard gates rather than quietly dropping to a smaller quant. NVIDIA L4 isn't a retail purchase for most people. It's rented by the hour. You can run this exact card on RunPod.

AI & Machine Learning benchmark results

Text Generation tok/s 143

SmolLM2-135M652.15
gemma-3-270m-it493.31
SmolLM2-360M440.21
Qwen1.5-0.5B431.41
Qwen2.5-0.5B-Instruct399.4
Qwen2.5-Coder-0.5B397.1
Qwen2 0.5B396.94
Qwen3 0.6B357.79
Qwen_Qwen3-0.6B356.93
LFM2.5-1.2B278.53
Llama 3.2 1B259.48
gemma-3-1b-it216.27
WorkloadResultTelemetryData
SmolLM2-135M652.15 tok/s
36 W63°CQ4_K_M
✓ Measured
gemma-3-270m-it493.31 tok/s
34 W51°CQ4_K_M
✓ Measured
SmolLM2-360M440.21 tok/s
34 W55°CQ4_K_M
✓ Measured
Qwen1.5-0.5B431.41 tok/s
41 W64°CQ4_K_M
✓ Measured
Qwen2 0.5B396.94 tok/s
38 W55°CQ4_K_M
✓ Measured
Qwen2.5-0.5B-Instruct399.4 tok/s
41 W53°CQ4_K_M
✓ Measured
Qwen2.5-Coder-0.5B397.1 tok/s
39 W63°CQ4_K_M
✓ Measured
Qwen3 0.6B357.79 tok/s
46 W51°CQ4_K_M
✓ Measured
Qwen_Qwen3-0.6B356.93 tok/s
40 W54°CQ4_K_M
✓ Measured
Llama 3.2 1B259.48 tok/s
48 W54°CQ4_K_M
✓ Measured
gemma-3-1b-it216.27 tok/s
50 W52°CQ4_K_M
✓ Measured
LFM2.5-1.2B278.53 tok/s
46 W58°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 1.5B191.38 tok/s
51 W53°CQ4_K_M
✓ Measured
Qwen2-1.5B189.18 tok/s
52 W64°CQ4_K_M
✓ Measured
Qwen2.5-1.5B-Instruct191.39 tok/s
53 W55°CQ4_K_M
✓ Measured
Qwen2.5-Coder-1.5B184.6 tok/s
47 W58°CQ4_K_M
✓ Measured
Qwen3 1.7B177.5 tok/s
51 W51°CQ4_K_M
✓ Measured
Qwen3-1.7B176.08 tok/s
52 W61°CQ4_K_M
✓ Measured
MiniCPM5 2B138.25 tok/s
52 W56°CQ4_K_M
✓ Measured
gemma-2-2b-it114.17 tok/s
57 W53°CQ4_K_M
✓ Measured
gemma-2-2b-it-abliterated113 tok/s
55 W63°CQ4_K_M
✓ Measured
LFM2.5 2.6B128.88 tok/s
53 W56°CQ4_K_M
✓ Measured
AI21-Jamba-Reasoning-3B107.59 tok/s
59 W64°CQ4_K_M
✓ Measured
Granite 4.1 3B96.18 tok/s
55 W66°CQ4_K_M
✓ Measured
Hermes-3-Llama-3.2-3B106.14 tok/s
56 W64°CQ4_K_M
✓ Measured
Llama 3.2 3B107.1 tok/s
55 W53°CQ4_K_M
✓ Measured
Llama-3.2-3B-Instruct-uncensored101.16 tok/s
56 W62°CQ4_K_M
✓ Measured
Nanbeige4.2-3B✕ Won't fit needs ~4 GBVRAM-gated at this precision✓ Measured
Qwen2.5-3B-Instruct110.84 tok/s
57 W55°CQ4_K_M
✓ Measured
Qwen2.5-Coder-3B109.42 tok/s
54 W58°CQ4_K_M
✓ Measured
SmolLM3 3B112.07 tok/s
57 W54°CQ4_K_M
✓ Measured
SmolLM3-3B110.77 tok/s
53 W61°CQ4_K_M
✓ Measured
Phi-4 Mini 3.8B88.48 tok/s
57 W54°CQ4_K_M
✓ Measured
Gemma 3 4B82.03 tok/s
60 W54°CQ4_K_M
✓ Measured
Nemotron 3 Nano 4B79.8 tok/s
58 W54°CQ4_K_M
✓ Measured
Qwen3-4B84 tok/s
59 W55°CQ4_K_M
✓ Measured
Qwen3-4B-Instruct-250783.13 tok/s
60 W65°CQ4_K_M
✓ Measured
Qwen3-4B-Instruct-250783.03 tok/s
60 W66°CQ4_K_M
✓ Measured
Qwen3-4B-Thinking-250783.12 tok/s
59 W63°CQ4_K_M
✓ Measured
phi-2114.99 tok/s
57 W63°CQ4_K_M
✓ Measured
Dolphin X1 Trinity Nano 6B166.6 tok/s
47 W62°CQ4_K_M
✓ Measured
Phi-3.5-mini-instruct90.56 tok/s
60 W51°CQ4_K_M
✓ Measured
Phi-4-mini87.21 tok/s
56 W63°CQ4_K_M
✓ Measured
DeepSeek Coder 7B Instruct v1.557.06 tok/s
59 W66°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 7B53.27 tok/s
63 W53°CQ4_K_M
✓ Measured
Llama-2-7B56.06 tok/s
63 W62°CQ4_K_M
✓ Measured
Mistral 7B v0.354.02 tok/s
64 W54°CQ4_K_M
✓ Measured
Mistral-7B-Instruct-v0.153.6 tok/s
63 W64°CQ4_K_M
✓ Measured
Mistral-7B-Instruct-v0.253.31 tok/s
63 W59°CQ4_K_M
✓ Measured
Mistral-7B-Instruct-v0.353.66 tok/s
63 W62°CQ4_K_M
✓ Measured
Qwen2.5-7B-Instruct53.32 tok/s
63 W55°CQ4_K_M
✓ Measured
Qwen2.5-Coder 7B53.26 tok/s
63 W55°CQ4_K_M
✓ Measured
Qwen2.5-Coder-7B-Instruct-abliterated53.02 tok/s
61 W64°CQ4_K_M
✓ Measured
Qwen2.5-VL 7B Instruct52.57 tok/s
54 W65°CQ4_K_M
✓ Measured
DarkIdol-Llama-3.1-8B-Instruct-1.2-Uncensored48.82 tok/s
60 W64°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill Llama 8B50.39 tok/s
62 W55°CQ4_K_M
✓ Measured
DeepSeek-R1-0528-Qwen3-8B49.07 tok/s
63 W52°CQ4_K_M
✓ Measured
Dolphin 3.0 Llama 3.1 8B50.24 tok/s
63 W64°CQ4_K_M
✓ Measured
Dolphin X1 8B50.12 tok/s
63 W62°CQ4_K_M
✓ Measured
Josiefied-Qwen3-8B-abliterated-v148.72 tok/s
62 W64°CQ4_K_M
✓ Measured
L3-8B-Stheno-v3.250.07 tok/s
61 W65°CQ4_K_M
✓ Measured
LFM2.5-8B-A1B170.2 tok/s
48 W63°CQ4_K_M
✓ Measured
Llama 3 8B50.35 tok/s
59 W67°CQ4_K_M
✓ Measured
Llama 3.1 8B50.45 tok/s
4.8 GB peak60 W50°C0.84 tok/WQ4_K_M
✓ Measured
Meta-Llama-3.1-8B-Instruct50.36 tok/s
62 W57°CQ4_K_M
✓ Measured
Qwen3 8B48.95 tok/s
63 W53°CQ4_K_M
✓ Measured
Qwen3-8B48.34 tok/s
60 W64°CQ4_K_M
✓ Measured
dolphin-2.9-llama3-8b50.04 tok/s
61 W66°CQ4_K_M
✓ Measured
Nemotron Nano 9B v235.8 tok/s
61 W67°CQ4_K_M
✓ Measured
Ornith 1.5 9B43.76 tok/s
61 W56°CQ4_K_M
✓ Measured
Ornith-1.0-9B43.35 tok/s
62 W53°CQ4_K_M
✓ Measured
gemma-2-9b33.8 tok/s
64 W64°CQ4_K_M
✓ Measured
Gemma 3 12B31.01 tok/s
65 W55°CQ4_K_M
✓ Measured
Gemma 3 12B (Q3_K_M)32.29 tok/s
62 W62°CQ3_K_M
✓ Measured
Gemma 4 12B31.49 tok/s
65 W56°CQ4_K_M
✓ Measured
NemoMix-Unleashed-12B33.02 tok/s
63 W64°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 14B27.39 tok/s
65 W59°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 14B (Q3_K_M)27.87 tok/s
65 W64°CQ3_K_M
✓ Measured
EVA-Qwen2.5-14B-v0.227.3 tok/s
64 W63°CQ4_K_M
✓ Measured
Hermes-4-14B27.48 tok/s
65 W65°CQ4_K_M
✓ Measured
Phi-4 14B27.24 tok/s
66 W57°CQ4_K_M
✓ Measured
Phi-4 14B (Q3_K_M)27.86 tok/s
64 W62°CQ3_K_M
✓ Measured
Qwen2.5-14B-Instruct27.37 tok/s
65 W55°CQ4_K_M
✓ Measured
Qwen2.5-Coder 14B27.46 tok/s
8.6 GB peak62 W57°C0.44 tok/WQ4_K_M
✓ Measured
Qwen2.5-Coder-14B-Instruct-abliterated27.29 tok/s
65 W65°CQ4_K_M
✓ Measured
Qwen3 14B27.61 tok/s
65 W55°CQ4_K_M
✓ Measured
Qwen3-14B27.48 tok/s
60 W65°CQ4_K_M
✓ Measured
Uncensored27.29 tok/s
65 W66°CQ4_K_M
✓ Measured
StarCoder2 15B24.4 tok/s
66 W60°CQ4_K_M
✓ Measured
Mistral-Nemo-Instruct-240733.01 tok/s
63 W64°CQ4_K_M
✓ Measured
gpt-oss-20b91.78 tok/s
54 W62°CQ4_K_M
✓ Measured
Codestral 22B18.14 tok/s
67 W61°CQ4_K_M
✓ Measured
Codestral 22B (Q3_K_M)17.96 tok/s
67 W68°CQ3_K_M
✓ Measured
GLM-4.7-Flash-REAP-23B-A3B70.87 tok/s
54 W65°CQ4_K_M
✓ Measured
DeepSeek-Coder-V2-Lite110.11 tok/s
51 W62°CQ4_K_M
✓ Measured
Cydonia-24B-v4.317.3 tok/s
66 W64°CQ4_K_M
✓ Measured
Devstral Small 24B17.34 tok/s
66 W60°CQ4_K_M
✓ Measured
Dolphin 3.0 R1 Mistral 24B17.32 tok/s
67 W63°CQ4_K_M
✓ Measured
Dolphin Mistral 24B Venice17.34 tok/s
67 W63°CQ4_K_M
✓ Measured
Dolphin-Mistral-24B-Venice-Edition17.2 tok/s
64 W64°CQ4_K_M
✓ Measured
Mistral Small 24B17.33 tok/s
66 W61°CQ4_K_M
✓ Measured
Mistral Small 24B (Q3_K_M)16.62 tok/s
67 W67°CQ3_K_M
✓ Measured
Gemma 4 26B A4B68.51 tok/s
54 W54°CQ4_K_M
✓ Measured
Gemma 3 27B14.18 tok/s
67 W63°CQ4_K_M
✓ Measured
Qwen3.6 27B14.2 tok/s
66 W58°CQ4_K_M
✓ Measured
Qwen3.8 27B13.91 tok/s
67 W58°CQ4_K_M
✓ Measured
Nemotron 3.5 Lightning 30B A3B✕ Won't fit needs ~30 GBVRAM-gated at this precision✓ Measured
Nemotron-3-Nano-30B-A3B✕ Won't fit needs ~31 GBVRAM-gated at this precision✓ Measured
Qwen3 30B A3B96.15 tok/s
52 W57°CQ4_K_M
✓ Measured
Qwen3 30B A3B (Q3_K_M)94.03 tok/s
52 W66°CQ3_K_M
✓ Measured
Qwen3 30B A3B Instruct 250799.74 tok/s
45 W70°CQ4_K_M
✓ Measured
Qwen3-30B-A3B95.02 tok/s
51 W63°CQ4_K_M
✓ Measured
Qwen3-Coder 30B A3B99.14 tok/s
53 W58°CQ4_K_M
✓ Measured
Gemma 4 31B12.84 tok/s
67 W61°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 32B12.28 tok/s
67 W63°CQ4_K_M
✓ Measured
DeepSeek-R1-Distill-Qwen-32B-abliterated12.26 tok/s
66 W66°CQ4_K_M
✓ Measured
Olmo-3.1-32B-Think12.27 tok/s
67 W63°CQ4_K_M
✓ Measured
QwQ 32B12.28 tok/s
67 W63°CQ4_K_M
✓ Measured
Qwen2.5-32B12.24 tok/s
65 W59°CQ4_K_M
✓ Measured
Qwen2.5-Coder 32B12.29 tok/s
67 W62°CQ4_K_M
✓ Measured
Qwen2.5-Coder 32B (Q3_K_M)12.05 tok/s
67 W68°CQ3_K_M
✓ Measured
Qwen3-32B12.42 tok/s
67 W65°CQ4_K_M
✓ Measured
Dolphin 2.9.1 Yi 1.5 34B11.92 tok/s
67 W65°CQ4_K_M
✓ Measured
Ornith 1.5 35B A3B84.69 tok/s
46 W68°CQ4_K_M
✓ Measured
Ornith-1.0-35B70.31 tok/s
52 W55°CQ4_K_M
✓ Measured
Qwen-AgentWorld-35B-A3B70.23 tok/s
53 W57°CQ4_K_M
✓ Measured
Qwen3.6 35B A3B78.98 tok/s
51 W55°CQ4_K_M
✓ Measured
GLM-4.7-Flash75.03 tok/s
52 W53°CQ4_K_M
✓ Measured
Laguna-XS-2.1✕ Won't fit needs ~26 GBVRAM-gated at this precision✓ Measured
Kwaipilot_KAT-Coder-V2.5-Dev79.01 tok/s
51 W54°CQ4_K_M
✓ Measured
DeepSeek-R1-Distill-Llama-70B✕ Won't fit needs ~54 GBVRAM-gated at this precision✓ Measured
Hermes-4-70B✕ Won't fit needs ~54 GBVRAM-gated at this precision✓ Measured
Llama 3.3 70B✕ Won't fit needs ~46 GBVRAM-gated at this precision✓ Measured
Llama-3.3-70B-Instruct-abliterated✕ Won't fit needs ~54 GBVRAM-gated at this precision✓ Measured
Meta-Llama-3.1-70B✕ Won't fit needs ~54 GBVRAM-gated at this precision✓ Measured
Qwen2.5-72B✕ Won't fit needs ~60 GBVRAM-gated at this precision✓ Measured
Qwen3-Next-80B-A3B✕ Won't fit needs ~61 GBVRAM-gated at this precision✓ Measured
Qwen3-Next-80B-A3B-Thinking✕ Won't fit needs ~61 GBVRAM-gated at this precision✓ Measured
Qwen3-Next-80B-A3B-Thinking✕ Won't fit needs ~61 GBVRAM-gated at this precision✓ Measured
Qwen3-Coder-Next✕ Won't fit needs ~61 GBVRAM-gated at this precision✓ Measured
Qwen3-Coder-Next✕ Won't fit needs ~61 GBVRAM-gated at this precision✓ Measured
Qwen3-Coder-Next-abliterated✕ Won't fit needs ~61 GBVRAM-gated at this precision✓ Measured
gpt-oss-120b✕ Won't fit needs ~70 GBVRAM-gated at this precision✓ Measured

Embeddings sentences/s 1

WorkloadResultTelemetryData
BGE-Large Embeddings1016.5 sentences/s
73 W58°C
✓ Measured

Image Generation images/min 15

SDXL Turbo229.86
Stable Diffusion 1.528.64
Sana 1.6B13.84
PixArt-Sigma XL8.08
Stable Diffusion XL5.18
Z-Image Turbo (1024px)4.6
Playground v2.53.01
Stable Diffusion 3.5 Medium2.96
Z-Image Turbo2.475
AuraFlow v0.31.02
Z-Image0.37
WorkloadResultTelemetryData
Stable Diffusion 1.528.64 images/min
72 W69°C
✓ Measured
SDXL Turbo229.86 images/min
46 W68°C
✓ Measured
Sana 1.6B13.84 images/min
72 W56°C
✓ Measured
Stable Diffusion XL5.18 images/min
14.6 GB peak72 W65°C11.6 s/img
✓ Measured
Playground v2.53.01 images/min
72 W80°C
✓ Measured
PixArt-Sigma XL8.08 images/min
72 W61°C
✓ Measured
Stable Diffusion 3.5 Medium2.96 images/min
72 W71°C
✓ Measured
Z-Image Turbo2.48 images/min
21.9 GB peak72 W74°C24.6 s/img
✓ Measured
Z-Image0.37 images/min
73 W78°C
✓ Measured
Z-Image Turbo (1024px)4.6 images/min
71 W81°C
✓ Measured
AuraFlow v0.31.02 images/min
72 W78°C
✓ Measured
FLUX.1 dev✕ Won't fit needs ~26 GBVRAM-gated at this precision✓ Measured
Stable Diffusion 3.5 Large✕ Won't fit needs ~30 GBVRAM-gated at this precision✓ Measured
FLUX.1 Schnell✕ Won't fit needs ~35 GBVRAM-gated at this precision✓ Measured
Krea 2 Turbo✕ Won't fit needs ~44 GBVRAM-gated at this precision✓ Measured

Image Editing images/min 2

WorkloadResultTelemetryData
FLUX.1 Kontext dev✕ Won't fit needs ~26 GBVRAM-gated at this precision✓ Measured
Qwen-Image-Edit✕ Won't fit needs ~42 GBVRAM-gated at this precision✓ Measured

Background Removal images/min 1

WorkloadResultTelemetryData
BiRefNet305.81 images/min
46 W50°C
✓ Measured

Upscaling images/min 1

WorkloadResultTelemetryData
Swin2SR 4x Upscaler18.23 images/min
72 W47°C
✓ Measured

Image to Video clips/min 5

WorkloadResultTelemetryData
Stable Video Diffusion0.63 clips/min
72 W79°C
✓ Measured
LTX-Video (image to video)1.41 clips/min
71 W82°C
✓ Measured
Wan 2.2 TI2V-5B (image to video)✕ Won't fit needs ~31 GBVRAM-gated at this precision✓ Measured
Stable Video Diffusion XT✕ Won't fit needs ~40 GBVRAM-gated at this precision✓ Measured
CogVideoX-5B I2V✕ Won't fit needs ~44 GBVRAM-gated at this precision✓ Measured

Video Generation frames/s 3

Wan 2.1 1.3B0.219
CogVideoX-2B0.179
Wan 2.2 5B (720p)0.15
WorkloadResultTelemetryData
Wan 2.2 5B (720p)0.15 frames/s
16.7 GB peak70 W84°C335 s/clip
✓ Measured
CPU offload
CogVideoX-2B0.18 frames/s
72 W79°C274 s/clip
✓ Measured
Wan 2.1 1.3B0.22 frames/s
72 W78°C223.9 s/clip
✓ Measured

Image to 3D assets/hour 4

TripoSR Image-to-3D967.5
TripoSG Image-to-3D83.2
TRELLIS Image-to-3D68.8
TRELLIS.2 Image-to-3D23.6
WorkloadResultTelemetryData
TripoSR Image-to-3D967.5 assets/hour
46 W71°C
✓ Measured
TripoSG Image-to-3D83.2 assets/hour
70 W80°C
✓ Measured
TRELLIS Image-to-3D68.8 assets/hour✓ Measured
TRELLIS.2 Image-to-3D23.6 assets/hour✓ Measured

Music Generation x realtime 3

ACE-Step 1.512.84
ACE-Step v1 3.5B8.26
MusicGen Small1.01
WorkloadResultTelemetryData
MusicGen Small1.01 x realtime
53 W57°C
✓ Measured
ACE-Step 1.512.84 x realtime
56 W76°C
✓ Measured
ACE-Step v1 3.5B8.26 x realtime
68 W64°C
✓ Measured

Sound Effects x realtime 3

MiDashengLM-Gen0.71
EzAudio XL0.38
MOSS-SoundEffect v2.00.37
WorkloadResultTelemetryData
EzAudio XL0.38 x realtime
72 W82°C
✓ Measured
MOSS-SoundEffect v2.00.37 x realtime
72 W79°C
✓ Measured
MiDashengLM-Gen0.71 x realtime
71 W79°C
✓ Measured

Speech to Text x realtime 1

WorkloadResultTelemetryData
Whisper large-v370.11 x realtime
56 W52°C
✓ Measured

Text to Speech x realtime 1

WorkloadResultTelemetryData
Kokoro TTS 82M97.35 x realtime
34 W46°C
✓ Measured

Depth Estimation images/min 2

WorkloadResultTelemetryData
Depth Anything V2 Small916.66 images/min
28 W47°C
✓ Measured
Depth Anything V2 Large770.37 images/min
30 W48°C
✓ Measured

Segmentation images/min 2

WorkloadResultTelemetryData
SAM ViT-Base315.89 images/min
40 W52°C
✓ Measured
SAM ViT-Huge70.46 images/min
69 W45°C
✓ Measured

Vision Language images/min 2

WorkloadResultTelemetryData
Florence-2 Base174.65 images/min
38 W41°C
✓ Measured
Florence-2 Large93.77 images/min
44 W43°C
✓ Measured

Fine-Tuning train tok/s 4

TinyLlama 1.1B LoRA4756
Qwen2.5 1.5B LoRA3720.9
SmolLM2 1.7B LoRA3240.2
Qwen2.5 7B LoRA1063.9
WorkloadResultTelemetryData
TinyLlama 1.1B LoRA4756 train tok/s
71 W42°C
✓ Measured
Qwen2.5 1.5B LoRA3720.9 train tok/s
71 W45°C
✓ Measured
SmolLM2 1.7B LoRA3240.2 train tok/s
72 W49°C
✓ Measured
Qwen2.5 7B LoRA1063.9 train tok/s
72 W57°C
✓ Measured

LLM Serving serve tok/s 4

TinyLlama 1.1B served2584.3
Qwen2.5 1.5B served1827
SmolLM2 1.7B served1524.9
Qwen2.5 7B served469.9
WorkloadResultTelemetryData
TinyLlama 1.1B served2584.3 serve tok/s
72 W59°C
✓ Measured
Qwen2.5 1.5B served1827 serve tok/s
72 W56°C
✓ Measured
SmolLM2 1.7B served1524.9 serve tok/s
72 W58°C
✓ Measured
Qwen2.5 7B served469.9 serve tok/s
72 W65°C
✓ Measured

Speculative Decoding x vs solo 1

WorkloadResultTelemetryData
Qwen2.5 1.5B + 0.5B draft0.97 x vs solo✓ Measured
How we measured this. Every result comes from our own pinned, reproducible AI suite, 12 workloads: the Qwen3-4B to Llama-70B LLM ladder (llama.cpp, Q4_K_M), SDXL / Z-Image / FLUX-dev generation, FLUX-Kontext / Qwen-Edit editing, and LTX / Wan video, run first-party on rented hardware with under 0.5% run-to-run variance. Peak VRAM, power draw, temperature and tokens-per-watt are captured per workload. “Won’t fit” rows are real data: where a model exceeds the card’s VRAM at the tested precision we record a hard gate rather than silently dropping to a smaller quant. Measured 2026-07-10 · harness 2.0.0.

NVIDIA L4 specifications

ArchitectureAda Lovelace
CUDA cores7,424
VRAM24GB GDDR6
Memory bus192-bit
Memory bandwidth300 GB/s
Boost clock2,040 MHz
TDP72 W
ProcessTSMC 4N
InterfacePCIe 4.0 x16
Release date2023-03-21
Launch MSRP$2,500

Verdict, capable, but 24GB sets the ceiling

NVIDIA L4 scores 5.0/100, #41 of 102. It ran 6 of 12; 5 exceeded its 24GB. Every figure here is our own measurement.

Relative performance: where the NVIDIA L4 lands

100% = this card, AI & Machine Learning headline metric (AI Score). #20 of 21 datacenter cards in this vertical.

GPURelative%AI Score
NVIDIA A40
354%17.7
NVIDIA A100 40GB SXM4
340%17
NVIDIA A100 40GB PCIe
334%16.7
NVIDIA A10G
128%6.4
NVIDIA L4
100%5
NVIDIA T4
58%2.9

← All AI & Machine Learning GPU rankings

The silicon

Transistors35,800 million
Die size294.5 mm²
Process node4 nm
Fabricated byTSMC
Transistor density121.6 million per mm²

Denser than 94% of the 746 cards we have silicon data for. Density is the clearest measure of what a process node bought: a card that gained it without growing the die got its speed from the fab rather than the architecture.

Silicon figures from Wikipedia (CC BY-SA 4.0). Benchmarks on this page are our own. Compare every chip.

What this card can build

Whole-job timings, composed from our measured per-model results on this card.

WorkflowTimeEnergyBasis
Depth pass on a batch17 s0.03 Whmeasured
Voiceovers from scripts36 s0.17 Whmeasured
Podcast episode pass2 min1.6 Whall 3 stages measured
Transcribe and subtitle videos4.4 min3.97 Whmeasured
Masking run7.5 min8.2 Whmeasured
Caption a training dataset10.8 min7.82 Whmeasured
Product catalogue cutout12.1 min13.61 Whall 2 stages measured
24-frame storyboard12.2 min13.69 Whall 2 stages measured
3D game asset kit22.6 min4.63 Whall 2 stages measured
Full codebase review36.4 min37.51 Whmeasured
Short social clips60.3 min69.46 Whall 3 stages measured
Photos to 3D models66.2 min0.06 Whall 2 stages measured

Can't run: Animate a batch of images (needs Wan 2.2 TI2V-5B (image to video)), Product shoot, start to finish (needs FLUX.1 Kontext dev), Product photo shoot (needs FLUX.1 Kontext dev), Photo restoration batch (needs FLUX.1 Kontext dev), Restore and enlarge photos (needs FLUX.1 Kontext dev), Character sheet, 12 poses (needs FLUX.1 dev), 6-panel comic page (needs FLUX.1 dev), Long-form article batch (needs Llama 3.3 70B).

Rent or buy?

This card is $2,500 to buy. The cheapest listed rate on Vast.ai is $0.312/hour, but that is the floor: we budget $0.374/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 6,677 GPU-hours. Below it you are paying for idle silicon.

How you would use itGPU-hours a yearRental cost a yearTime to break even
2 hours a day, hobby730$2739.1 years
8 hours a day, working on it2,920$1,0932.3 years
24/7, always-on agent8,760$3,2809.1 months

At hobby usage this card is very unlikely to pay for itself before it is superseded. Rent it. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.

Rental price

$0.440/hr+12.8% since 2026-08-14low $0.269 · high $0.440

Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.