NVIDIA T4, AI & Machine Learning Benchmarks & Specs

16GB · AI Score 2.9/100 · first-party measured on 12 AI workloads

2.9 AI Score ✓ Measured

Every number on this page is first-party: NVIDIA T4 was run on our pinned 12-workload AI suite on 2026-07-10, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) NVIDIA T4 delivers about 35.01 tokens/sec. Llama 3.3 70B does not fit. It needs roughly 42GB and this card has 16GB. For image generation, SDXL runs at 1.18 it/s, while FLUX.1-dev won't fit at BF16 (needs ~26GB). 6 of the 12 workloads won't fit on 16GB at the tested precision, Qwen3 32B, Llama 3.3 70B, FLUX.1-dev, FLUX.1 Kontext and others. We publish those as hard gates rather than quietly dropping to a smaller quant.

AI & Machine Learning benchmark results

Text Generation tok/s 123

SmolLM2-135M416.24
gemma-3-270m362.8
Qwen1.5-0.5B328.29
Qwen2.5-0.5B295.47
Qwen2.5-Coder-0.5B294.96
Qwen3 0.6B263.47
Qwen3-0.6B259.91
LFM2.5-1.2B221.66
Llama 3.2 1B207.71
gemma-3-1b156.88
DeepSeek-R1 Distill 1.5B148.72
Qwen2.5-Coder-1.5B148.51
WorkloadResultTelemetryData
SmolLM2-135M416.24 tok/s
36 W50°CQ4_K_M
✓ Measured
gemma-3-270m362.8 tok/s
37 W45°CQ4_K_M
✓ Measured
Qwen1.5-0.5B328.29 tok/s
42 W49°CQ4_K_M
✓ Measured
Qwen2.5-0.5B295.47 tok/s
43 W44°CQ4_K_M
✓ Measured
Qwen2.5-Coder-0.5B294.96 tok/s
44 W49°CQ4_K_M
✓ Measured
Qwen3 0.6B263.47 tok/s
47 W31°CQ4_K_M
✓ Measured
Qwen3-0.6B259.91 tok/s
49 W45°CQ4_K_M
✓ Measured
LFM2.5-1.2B221.66 tok/s
47 W46°CQ4_K_M
✓ Measured
Llama 3.2 1B207.71 tok/s
48 W32°CQ4_K_M
✓ Measured
gemma-3-1b156.88 tok/s
51 W45°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 1.5B148.72 tok/s
49 W47°CQ4_K_M
✓ Measured
Qwen2.5-Coder-1.5B148.51 tok/s
52 W45°CQ4_K_M
✓ Measured
Qwen2.5-1.5B148.14 tok/s
52 W44°CQ4_K_M
✓ Measured
Qwen2-1.5B147.68 tok/s
55 W51°CQ4_K_M
✓ Measured
Qwen3 1.7B139.92 tok/s
57 W34°CQ4_K_M
✓ Measured
LFM2.5-8B-A1B139.03 tok/s
55 W51°CQ4_K_M
✓ Measured
Qwen3-1.7B135.21 tok/s
47 W48°CQ4_K_M
✓ Measured
Dolphin X1 Trinity Nano 6B114.21 tok/s
54 W49°CQ4_K_M
✓ Measured
gemma-2-2b92.13 tok/s
53 W46°CQ4_K_M
✓ Measured
gemma-2-2b-it-abliterated91.91 tok/s
53 W46°CQ4_K_M
✓ Measured
phi-289.6 tok/s
57 W51°CQ4_K_M
✓ Measured
AI21-Jamba-Reasoning-3B86.91 tok/s
62 W51°CQ4_K_M
✓ Measured
SmolLM3-3B86.86 tok/s
55 W48°CQ4_K_M
✓ Measured
Llama-3.2-3B-Instruct-uncensored86.29 tok/s
56 W48°CQ4_K_M
✓ Measured
SmolLM3 3B86.08 tok/s
58 W49°CQ4_K_M
✓ Measured
Qwen2.5-3B86.04 tok/s
55 W45°CQ4_K_M
✓ Measured
Llama 3.2 3B85.94 tok/s
54 W47°CQ4_K_M
✓ Measured
Qwen2.5-Coder-3B85.74 tok/s
59 W45°CQ4_K_M
✓ Measured
Hermes-3-Llama-3.2-3B84.99 tok/s
61 W50°CQ4_K_M
✓ Measured
DeepSeek-Coder-V2-Lite83.49 tok/s
55 W44°CQ4_K_M
✓ Measured
Phi-3.5-mini72.07 tok/s
52 W44°CQ4_K_M
✓ Measured
Qwen3-4B67.71 tok/s
53 W45°CQ4_K_M
✓ Measured
Qwen3-4B-Instruct-250767.45 tok/s
60 W51°CQ4_K_M
✓ Measured
Qwen3-4B-Thinking-250767 tok/s
62 W52°CQ4_K_M
✓ Measured
Qwen3-4B-Instruct-250766.89 tok/s
61 W52°CQ4_K_M
✓ Measured
Phi-4-mini66 tok/s
59 W46°CQ4_K_M
✓ Measured
Gemma 3 4B65.12 tok/s
60 W47°CQ4_K_M
✓ Measured
Phi-4 Mini 3.8B64.15 tok/s
59 W48°CQ4_K_M
✓ Measured
gpt-oss-20b63.63 tok/s
52 W45°CQ4_K_M
✓ Measured
Qwen3 30B A3B (Q3_K_M)61.82 tok/s
56 W49°CQ3_K_M
✓ Measured
GLM-4.7-Flash-REAP-23B-A3B53.83 tok/s
56 W51°CQ4_K_M
✓ Measured
Llama-2-7B42.98 tok/s
62 W45°CQ4_K_M
✓ Measured
Mistral-7B-Instruct-v0.241.31 tok/s
63 W45°CQ4_K_M
✓ Measured
Mistral-7B-Instruct-v0.140.63 tok/s
59 W48°CQ4_K_M
✓ Measured
Mistral 7B v0.340.01 tok/s
64 W49°CQ4_K_M
✓ Measured
Mistral-7B-Instruct-v0.339.87 tok/s
62 W49°CQ4_K_M
✓ Measured
Qwen2.5-7B39.01 tok/s
64 W46°CQ4_K_M
✓ Measured
Meta-Llama-3.1-8B37.8 tok/s
61 W45°CQ4_K_M
✓ Measured
DeepSeek-R1-0528-Qwen3-8B37.7 tok/s
58 W45°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 7B37.58 tok/s
63 W48°CQ4_K_M
✓ Measured
Qwen2.5-Coder 7B37.54 tok/s
63 W49°CQ4_K_M
✓ Measured
Qwen3-8B37.37 tok/s
62 W46°CQ4_K_M
✓ Measured
Qwen2.5-Coder-7B-Instruct-abliterated37.31 tok/s
63 W51°CQ4_K_M
✓ Measured
Qwen3 8B36.51 tok/s
61 W48°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill Llama 8B36.36 tok/s
63 W49°CQ4_K_M
✓ Measured
L3-8B-Stheno-v3.236.28 tok/s
55 W50°CQ4_K_M
✓ Measured
DarkIdol-Llama-3.1-8B-Instruct-1.2-Uncensored35.73 tok/s
62 W51°CQ4_K_M
✓ Measured
Dolphin 3.0 Llama 3.1 8B35.68 tok/s
61 W50°CQ4_K_M
✓ Measured
Dolphin X1 8B35.66 tok/s
61 W50°CQ4_K_M
✓ Measured
Josiefied-Qwen3-8B-abliterated-v135.62 tok/s
62 W51°CQ4_K_M
✓ Measured
dolphin-2.9-llama3-8b35.1 tok/s
62 W52°CQ4_K_M
✓ Measured
Llama 3.1 8B35.01 tok/s
4.7 GB peak60 W50°C0.59 tok/WQ4_K_M
✓ Measured
Ornith-1.0-9B32.93 tok/s
61 W46°CQ4_K_M
✓ Measured
gemma-2-9b28.62 tok/s
62 W48°CQ4_K_M
✓ Measured
Gemma 4 12B24.03 tok/s
62 W50°CQ4_K_M
✓ Measured
NemoMix-Unleashed-12B23.19 tok/s
63 W53°CQ4_K_M
✓ Measured
Gemma 3 12B23.12 tok/s
63 W49°CQ4_K_M
✓ Measured
Mistral-Nemo-Instruct-240722.98 tok/s
64 W52°CQ4_K_M
✓ Measured
EVA-Qwen2.5-14B-v0.219.92 tok/s
64 W47°CQ4_K_M
✓ Measured
Qwen2.5-14B19.84 tok/s
63 W48°CQ4_K_M
✓ Measured
Gemma 3 12B (Q3_K_M)19.79 tok/s
64 W52°CQ3_K_M
✓ Measured
Qwen2.5-Coder-14B19.76 tok/s
61 W48°CQ4_K_M
✓ Measured
Qwen3-14B19.44 tok/s
63 W50°CQ4_K_M
✓ Measured
Qwen3 14B19.31 tok/s
64 W50°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 14B19.21 tok/s
63 W50°CQ4_K_M
✓ Measured
Hermes-4-14B19.02 tok/s
64 W53°CQ4_K_M
✓ Measured
Uncensored18.94 tok/s
65 W53°CQ4_K_M
✓ Measured
Qwen2.5-Coder-14B-Instruct-abliterated18.86 tok/s
65 W52°CQ4_K_M
✓ Measured
Phi-4 14B17.46 tok/s
63 W50°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 14B (Q3_K_M)16.86 tok/s
64 W53°CQ3_K_M
✓ Measured
Phi-4 14B (Q3_K_M)16.04 tok/s
65 W52°CQ3_K_M
✓ Measured
StarCoder2 15B15.66 tok/s
65 W49°CQ4_K_M
✓ Measured
Codestral 22B12.59 tok/s
65 W50°CQ4_K_M
✓ Measured
Devstral Small 24B11.4 tok/s
64 W50°CQ4_K_M
✓ Measured
Dolphin Mistral 24B Venice11.36 tok/s
65 W51°CQ4_K_M
✓ Measured
Mistral Small 24B11.33 tok/s
64 W51°CQ4_K_M
✓ Measured
Dolphin-Mistral-24B-Venice-Edition11.33 tok/s
62 W52°CQ4_K_M
✓ Measured
Dolphin 3.0 R1 Mistral 24B11.32 tok/s
63 W51°CQ4_K_M
✓ Measured
Codestral 22B (Q3_K_M)11.3 tok/s
66 W50°CQ3_K_M
✓ Measured
Cydonia-24B-v4.311.21 tok/s
63 W53°CQ4_K_M
✓ Measured
Mistral Small 24B (Q3_K_M)10.13 tok/s
66 W54°CQ3_K_M
✓ Measured
Qwen3 32B✕ Won't fit needs ~23 GBVRAM-gated at this precision✓ Measured
Llama 3.3 70B✕ Won't fit needs ~46 GBVRAM-gated at this precision✓ Measured
Olmo-3.1-32B-Think✕ Won't fit needs ~25 GBVRAM-gated at this precision✓ Measured
DeepSeek-R1-Distill-Llama-70B✕ Won't fit needs ~54 GBVRAM-gated at this precision✓ Measured
DeepSeek-R1-Distill-Qwen-32B-abliterated✕ Won't fit needs ~25 GBVRAM-gated at this precision✓ Measured
DeepSeek-R1 Distill 32B✕ Won't fit needs ~21 GBVRAM-gated at this precision✓ Measured
Dolphin 2.9.1 Yi 1.5 34B✕ Won't fit needs ~21 GBVRAM-gated at this precision✓ Measured
Gemma 3 27B✕ Won't fit needs ~18 GBVRAM-gated at this precision✓ Measured
GLM-4.7-Flash✕ Won't fit needs ~23 GBVRAM-gated at this precision✓ Measured
Hermes-4-70B✕ Won't fit needs ~54 GBVRAM-gated at this precision✓ Measured
Qwen3-Coder-Next-abliterated✕ Won't fit needs ~61 GBVRAM-gated at this precision✓ Measured
KAT-Coder-V2.5-Dev✕ Won't fit needs ~27 GBVRAM-gated at this precision✓ Measured
Laguna-XS-2.1✕ Won't fit needs ~26 GBVRAM-gated at this precision✓ Measured
Llama-3.3-70B-Instruct-abliterated✕ Won't fit needs ~54 GBVRAM-gated at this precision✓ Measured
Meta-Llama-3.1-70B✕ Won't fit needs ~54 GBVRAM-gated at this precision✓ Measured
Nanbeige4.2-3B✕ Won't fit needs ~4 GBVRAM-gated at this precision✓ Measured
Nemotron-3-Nano-30B-A3B✕ Won't fit needs ~31 GBVRAM-gated at this precision✓ Measured
Ornith-1.0-35B✕ Won't fit needs ~28 GBVRAM-gated at this precision✓ Measured
Qwen-AgentWorld-35B-A3B✕ Won't fit needs ~28 GBVRAM-gated at this precision✓ Measured
Qwen3-30B-A3B✕ Won't fit needs ~24 GBVRAM-gated at this precision✓ Measured
Qwen3-Coder-Next✕ Won't fit needs ~61 GBVRAM-gated at this precision✓ Measured
Qwen3-Next-80B-A3B-Thinking✕ Won't fit needs ~61 GBVRAM-gated at this precision✓ Measured
Qwen2.5-32B✕ Won't fit needs ~25 GBVRAM-gated at this precision✓ Measured
Qwen2.5-72B✕ Won't fit needs ~60 GBVRAM-gated at this precision✓ Measured
Qwen2.5-Coder 32B (Q3_K_M)✕ Won't fit needs ~16 GBVRAM-gated at this precision✓ Measured
Qwen2.5-Coder 32B✕ Won't fit needs ~21 GBVRAM-gated at this precision✓ Measured
Qwen3 30B A3B✕ Won't fit needs ~20 GBVRAM-gated at this precision✓ Measured
Qwen3-Coder 30B A3B✕ Won't fit needs ~20 GBVRAM-gated at this precision✓ Measured
Qwen3-Coder-Next✕ Won't fit needs ~61 GBVRAM-gated at this precision✓ Measured
Qwen3-Next-80B-A3B-Thinking✕ Won't fit needs ~61 GBVRAM-gated at this precision✓ Measured
Qwen3-Next-80B-A3B✕ Won't fit needs ~61 GBVRAM-gated at this precision✓ Measured
QwQ 32B✕ Won't fit needs ~21 GBVRAM-gated at this precision✓ Measured

Image Generation images/min 8

WorkloadResultTelemetryData
Stable Diffusion XL2.36 images/min
12.7 GB peak69 W58°C25.5 s/img
✓ Measured
Sana 1.6B1.13 images/min
69 W46°C
✓ Measured
FLUX.1 dev✕ Won't fit needs ~26 GBVRAM-gated at this precision✓ Measured
FLUX.1 Schnell✕ Won't fit needs ~26 GBVRAM-gated at this precision✓ Measured
AuraFlow v0.3✕ Won't fit needs ~24 GBVRAM-gated at this precision✓ Measured
PixArt-Sigma XL✕ Won't fit needs ~12 GBVRAM-gated at this precision✓ Measured
Stable Diffusion 3.5 Large✕ Won't fit needs ~24 GBVRAM-gated at this precision✓ Measured
Stable Diffusion 3.5 Medium✕ Won't fit needs ~12 GBVRAM-gated at this precision✓ Measured

Fine-Tuning train tok/s 4

SmolLM2 1.7B LoRA1561.9
TinyLlama 1.1B LoRA1520.1
Qwen2.5 1.5B LoRA1434.4
WorkloadResultTelemetryData
SmolLM2 1.7B LoRA1561.9 train tok/s
69 W50°C
✓ Measured
TinyLlama 1.1B LoRA1520.1 train tok/s
68 W45°C
✓ Measured
Qwen2.5 1.5B LoRA1434.4 train tok/s
67 W48°C
✓ Measured
Qwen2.5 7B LoRA✕ Won't fit needs ~20 GBVRAM-gated at this precision✓ Measured

LLM Serving serve tok/s 4

TinyLlama 1.1B served1897.8
Qwen2.5 1.5B served1337.1
SmolLM2 1.7B served1222.5
WorkloadResultTelemetryData
TinyLlama 1.1B served1897.8 serve tok/s
72 W33°C
✓ Measured
Qwen2.5 1.5B served1337.1 serve tok/s
70 W36°C
✓ Measured
SmolLM2 1.7B served1222.5 serve tok/s
70 W37°C
✓ Measured
Qwen2.5 7B served✕ Won't fit needs ~20 GBVRAM-gated at this precision✓ Measured

Image Editing images/min 2

WorkloadResultTelemetryData
FLUX.1 Kontext dev✕ Won't fit needs ~26 GBVRAM-gated at this precision✓ Measured
Qwen-Image-Edit✕ Won't fit needs ~42 GBVRAM-gated at this precision✓ Measured

Depth Estimation images/min 2

WorkloadResultTelemetryData
Depth Anything V2 Small612.25 images/min
27 W28°C
✓ Measured
Depth Anything V2 Large402.34 images/min
49 W29°C
✓ Measured

Segmentation images/min 2

WorkloadResultTelemetryData
SAM ViT-Base180.73 images/min
67 W31°C
✓ Measured
SAM ViT-Huge34.91 images/min
70 W35°C
✓ Measured

Vision Language images/min 2

WorkloadResultTelemetryData
Florence-2 Base148.61 images/min
35 W26°C
✓ Measured
Florence-2 Large83.43 images/min
49 W28°C
✓ Measured

Speculative Decoding x vs solo 2

WorkloadResultTelemetryData
Qwen2.5 1.5B + 0.5B draft0.9 x vs solo✓ Measured
Qwen2.5 7B + 0.5B draft✕ Won't fit needs ~24 GBVRAM-gated at this precision✓ Measured

Video Generation frames/s 1

WorkloadResultTelemetryData
Wan 2.2 5B (720p)✕ Won't fit needs ~18 GBVRAM-gated at this precision✓ Measured

Background Removal images/min 1

WorkloadResultTelemetryData
BiRefNet183.48 images/min
52 W30°C
✓ Measured

Upscaling images/min 1

WorkloadResultTelemetryData
Swin2SR 4x Upscaler10.53 images/min
65 W39°C
✓ Measured

Text to Speech x realtime 1

WorkloadResultTelemetryData
Kokoro TTS 82M42.88 x realtime
51 W38°C
✓ Measured

Speech to Text x realtime 1

WorkloadResultTelemetryData
Whisper large-v344.36 x realtime
65 W40°C
✓ Measured

Music Generation x realtime 1

WorkloadResultTelemetryData
MusicGen Small0.92 x realtime
58 W47°C
✓ Measured
How we measured this. Every result comes from our own pinned, reproducible AI suite, 12 workloads: the Qwen3-4B to Llama-70B LLM ladder (llama.cpp, Q4_K_M), SDXL / Z-Image / FLUX-dev generation, FLUX-Kontext / Qwen-Edit editing, and LTX / Wan video, run first-party on rented hardware with under 0.5% run-to-run variance. Peak VRAM, power draw, temperature and tokens-per-watt are captured per workload. “Won’t fit” rows are real data: where a model exceeds the card’s VRAM at the tested precision we record a hard gate rather than silently dropping to a smaller quant. Measured 2026-07-10 · harness 2.0.0.

NVIDIA T4 specifications

ArchitectureTuring
CUDA cores2,560
VRAM16GB GDDR6
Memory bus256-bit
Memory bandwidth320 GB/s
Boost clock1,590 MHz
TDP70 W
Process12nm
InterfacePCIe 3.0 x16
Release date2018-09-13
Launch MSRP$2,299

Verdict, capable, but 16GB sets the ceiling

NVIDIA T4 scores 2.9/100, #70 of 102. It ran 4 of 12; 6 exceeded its 16GB. Every figure here is our own measurement.

Relative performance: where the NVIDIA T4 lands

100% = this card, AI & Machine Learning headline metric (AI Score). #21 of 21 desktop cards in this vertical.

GPURelative%AI Score
NVIDIA A100 40GB SXM4
586%17
NVIDIA A100 40GB PCIe
576%16.7
NVIDIA A10G
221%6.4
NVIDIA L4
172%5
NVIDIA T4
100%2.9

← All AI & Machine Learning GPU rankings

What this card can build

Whole-job timings, composed from our measured per-model results on this card.

WorkflowTimeEnergyBasis
50-image depth pass13 s0.1 Whmeasured
500-image masking run14.4 min16.59 Whmeasured
200-product catalogue cutout20.3 min21.65 Whall 2 stages measured
Full codebase review50.7 min51.37 Whmeasured

Can't run: 60-second AI short film (needs Qwen3 32B), 60-second AI short film, narrated (needs Qwen3 32B), 30-minute podcast pass (needs Qwen3 32B), 10 short social clips (needs Qwen3 32B), 40-product photo shoot (needs FLUX.1 Kontext dev), 6-panel comic page (needs Qwen3 32B), 20 long-form articles (needs Llama 3.3 70B), Character sheet, 12 poses (needs FLUX.1 dev), 100-photo restoration batch (needs FLUX.1 Kontext dev), 24-frame storyboard (needs Qwen3 32B), 100-photo restore and enlarge (needs FLUX.1 Kontext dev), 40-product shoot, start to finish (needs FLUX.1 Kontext dev).