NVIDIA A10G, AI & Machine Learning Benchmarks & Specs

24GB · AI Score 6.4/100 · first-party measured on 12 AI workloads

6.4 AI Score ✓ Measured

Every number on this page is first-party: NVIDIA A10G was run on our pinned 12-workload AI suite on 2026-07-10, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) NVIDIA A10G delivers about 86.6 tokens/sec. Llama 3.3 70B does not fit. It needs roughly 42GB and this card has 24GB. For image generation, SDXL runs at 3.19 it/s, while FLUX.1-dev won't fit at BF16 (needs ~26GB). 5 of the 12 workloads won't fit on 24GB at the tested precision, Qwen3 32B, Llama 3.3 70B, FLUX.1-dev, FLUX.1 Kontext and others. We publish those as hard gates rather than quietly dropping to a smaller quant.

AI & Machine Learning benchmark results

Text Generation tok/s 123

SmolLM2-135M680.37
gemma-3-270m597.49
Qwen1.5-0.5B542.19
Qwen2.5-Coder-0.5B513.36
Qwen2.5-0.5B508.78
Qwen3 0.6B451.38
Qwen3-0.6B446.74
LFM2.5-1.2B429.96
Llama 3.2 1B400.6
gemma-3-1b284.78
LFM2.5-8B-A1B268.04
Qwen3 1.7B267.08
WorkloadResultTelemetryData
SmolLM2-135M680.37 tok/s
71 W56°CQ4_K_M
✓ Measured
gemma-3-270m597.49 tok/s
68 W54°CQ4_K_M
✓ Measured
Qwen1.5-0.5B542.19 tok/s
74 W54°CQ4_K_M
✓ Measured
Qwen2.5-Coder-0.5B513.36 tok/s
79 W54°CQ4_K_M
✓ Measured
Qwen2.5-0.5B508.78 tok/s
76 W51°CQ4_K_M
✓ Measured
Qwen3 0.6B451.38 tok/s
75 W34°CQ4_K_M
✓ Measured
Qwen3-0.6B446.74 tok/s
80 W55°CQ4_K_M
✓ Measured
LFM2.5-1.2B429.96 tok/s
82 W58°CQ4_K_M
✓ Measured
Llama 3.2 1B400.6 tok/s
95 W46°CQ4_K_M
✓ Measured
gemma-3-1b284.78 tok/s
97 W55°CQ4_K_M
✓ Measured
LFM2.5-8B-A1B268.04 tok/s
96 W58°CQ4_K_M
✓ Measured
Qwen3 1.7B267.08 tok/s
100 W36°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 1.5B265.75 tok/s
102 W53°CQ4_K_M
✓ Measured
Qwen2.5-1.5B265.7 tok/s
81 W54°CQ4_K_M
✓ Measured
Qwen2-1.5B265.45 tok/s
96 W56°CQ4_K_M
✓ Measured
Qwen2.5-Coder-1.5B264.79 tok/s
94 W57°CQ4_K_M
✓ Measured
Qwen3-1.7B262.49 tok/s
101 W59°CQ4_K_M
✓ Measured
gemma-2-2b175.03 tok/s
114 W56°CQ4_K_M
✓ Measured
gemma-2-2b-it-abliterated174.56 tok/s
110 W59°CQ4_K_M
✓ Measured
phi-2172.45 tok/s
112 W58°CQ4_K_M
✓ Measured
SmolLM3 3B172.15 tok/s
107 W56°CQ4_K_M
✓ Measured
SmolLM3-3B171.66 tok/s
111 W59°CQ4_K_M
✓ Measured
Dolphin X1 Trinity Nano 6B171.23 tok/s
93 W60°CQ4_K_M
✓ Measured
Llama 3.2 3B170.47 tok/s
107 W47°CQ4_K_M
✓ Measured
Hermes-3-Llama-3.2-3B169.23 tok/s
102 W55°CQ4_K_M
✓ Measured
Qwen2.5-3B168.23 tok/s
111 W56°CQ4_K_M
✓ Measured
Qwen2.5-Coder-3B168 tok/s
108 W57°CQ4_K_M
✓ Measured
Llama-3.2-3B-Instruct-uncensored167.63 tok/s
104 W58°CQ4_K_M
✓ Measured
AI21-Jamba-Reasoning-3B166.29 tok/s
107 W55°CQ4_K_M
✓ Measured
DeepSeek-Coder-V2-Lite160.13 tok/s
104 W57°CQ4_K_M
✓ Measured
Phi-3.5-mini145.97 tok/s
120 W53°CQ4_K_M
✓ Measured
gpt-oss-20b144.61 tok/s
99 W60°CQ4_K_M
✓ Measured
Phi-4 Mini 3.8B144.42 tok/s
113 W52°CQ4_K_M
✓ Measured
Phi-4-mini144.27 tok/s
114 W60°CQ4_K_M
✓ Measured
Qwen3-Coder 30B A3B142.77 tok/s
97 W59°CQ4_K_M
✓ Measured
Qwen3-30B-A3B140.14 tok/s
97 W58°CQ4_K_M
✓ Measured
Qwen3 30B A3B140.12 tok/s
101 W59°CQ4_K_M
✓ Measured
Qwen3-4B-Instruct-2507130.83 tok/s
111 W56°CQ4_K_M
✓ Measured
Qwen3-4B-Thinking-2507129.99 tok/s
117 W58°CQ4_K_M
✓ Measured
Qwen3-4B-Instruct-2507129.9 tok/s
118 W59°CQ4_K_M
✓ Measured
Qwen3-4B129.87 tok/s
112 W57°CQ4_K_M
✓ Measured
Gemma 3 4B124.85 tok/s
116 W49°CQ4_K_M
✓ Measured
Qwen3 30B A3B (Q3_K_M)120.38 tok/s
102 W60°CQ3_K_M
✓ Measured
KAT-Coder-V2.5-Dev113.6 tok/s
87 W55°CQ4_K_M
✓ Measured
Ornith-1.0-35B104.9 tok/s
97 W56°CQ4_K_M
✓ Measured
Qwen-AgentWorld-35B-A3B104.36 tok/s
96 W56°CQ4_K_M
✓ Measured
GLM-4.7-Flash103.24 tok/s
101 W53°CQ4_K_M
✓ Measured
GLM-4.7-Flash-REAP-23B-A3B97.48 tok/s
110 W58°CQ4_K_M
✓ Measured
Llama-2-7B96.69 tok/s
121 W58°CQ4_K_M
✓ Measured
Mistral 7B v0.392.82 tok/s
125 W56°CQ4_K_M
✓ Measured
Mistral-7B-Instruct-v0.292.74 tok/s
124 W57°CQ4_K_M
✓ Measured
Mistral-7B-Instruct-v0.192.45 tok/s
128 W60°CQ4_K_M
✓ Measured
Mistral-7B-Instruct-v0.392.45 tok/s
119 W59°CQ4_K_M
✓ Measured
Qwen2.5-Coder-7B-Instruct-abliterated91.41 tok/s
110 W55°CQ4_K_M
✓ Measured
Qwen2.5-Coder 7B90.05 tok/s
121 W56°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 7B89.79 tok/s
122 W54°CQ4_K_M
✓ Measured
Qwen2.5-7B89.68 tok/s
120 W58°CQ4_K_M
✓ Measured
dolphin-2.9-llama3-8b87.98 tok/s
114 W56°CQ4_K_M
✓ Measured
DarkIdol-Llama-3.1-8B-Instruct-1.2-Uncensored87.84 tok/s
113 W56°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill Llama 8B86.83 tok/s
122 W55°CQ4_K_M
✓ Measured
L3-8B-Stheno-v3.286.8 tok/s
118 W58°CQ4_K_M
✓ Measured
Meta-Llama-3.1-8B86.69 tok/s
118 W57°CQ4_K_M
✓ Measured
Llama-3.1-8B86.62 tok/s
120 W58°CQ4_K_M
✓ Measured
Dolphin X1 8B86.61 tok/s
121 W61°CQ4_K_M
✓ Measured
Dolphin 3.0 Llama 3.1 8B86.54 tok/s
122 W61°CQ4_K_M
✓ Measured
Qwen3 8B84.09 tok/s
117 W41°CQ4_K_M
✓ Measured
Josiefied-Qwen3-8B-abliterated-v183.26 tok/s
121 W59°CQ4_K_M
✓ Measured
DeepSeek-R1-0528-Qwen3-8B83.22 tok/s
122 W54°CQ4_K_M
✓ Measured
Qwen3-8B83.04 tok/s
120 W60°CQ4_K_M
✓ Measured
Ornith-1.0-9B73.78 tok/s
121 W55°CQ4_K_M
✓ Measured
NemoMix-Unleashed-12B57.52 tok/s
118 W56°CQ4_K_M
✓ Measured
gemma-2-9b57.37 tok/s
128 W60°CQ4_K_M
✓ Measured
Mistral-Nemo-Instruct-240757 tok/s
125 W59°CQ4_K_M
✓ Measured
Gemma 4 12B52.67 tok/s
125 W58°CQ4_K_M
✓ Measured
Gemma 3 12B52.07 tok/s
128 W52°CQ4_K_M
✓ Measured
Phi-4 14B48.18 tok/s
128 W59°CQ4_K_M
✓ Measured
Qwen3 14B48.12 tok/s
126 W46°CQ4_K_M
✓ Measured
Qwen3-14B47.8 tok/s
124 W60°CQ4_K_M
✓ Measured
Hermes-4-14B47.8 tok/s
130 W60°CQ4_K_M
✓ Measured
Uncensored47.45 tok/s
124 W57°CQ4_K_M
✓ Measured
Qwen2.5-Coder-14B47.07 tok/s
123 W57°CQ4_K_M
✓ Measured
Qwen2.5-14B47.02 tok/s
125 W56°CQ4_K_M
✓ Measured
EVA-Qwen2.5-14B-v0.247 tok/s
130 W58°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 14B46.96 tok/s
131 W59°CQ4_K_M
✓ Measured
Qwen2.5-Coder-14B-Instruct-abliterated46.91 tok/s
127 W59°CQ4_K_M
✓ Measured
Gemma 3 12B (Q3_K_M)46.29 tok/s
126 W57°CQ3_K_M
✓ Measured
Phi-4 14B (Q3_K_M)43.56 tok/s
126 W57°CQ3_K_M
✓ Measured
StarCoder2 15B42.19 tok/s
132 W62°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 14B (Q3_K_M)41.72 tok/s
126 W58°CQ3_K_M
✓ Measured
Codestral 22B32.09 tok/s
133 W61°CQ4_K_M
✓ Measured
Cydonia-24B-v4.331.04 tok/s
131 W58°CQ4_K_M
✓ Measured
Dolphin-Mistral-24B-Venice-Edition31 tok/s
129 W59°CQ4_K_M
✓ Measured
Devstral Small 24B30.98 tok/s
133 W60°CQ4_K_M
✓ Measured
Dolphin Mistral 24B Venice30.95 tok/s
131 W61°CQ4_K_M
✓ Measured
Mistral Small 24B30.94 tok/s
132 W61°CQ4_K_M
✓ Measured
Dolphin 3.0 R1 Mistral 24B30.9 tok/s
131 W62°CQ4_K_M
✓ Measured
Codestral 22B (Q3_K_M)28.01 tok/s
133 W62°CQ3_K_M
✓ Measured
Mistral Small 24B (Q3_K_M)26.8 tok/s
131 W59°CQ3_K_M
✓ Measured
Gemma 3 27B24.8 tok/s
133 W63°CQ4_K_M
✓ Measured
DeepSeek-R1-Distill-Qwen-32B-abliterated22.14 tok/s
127 W58°CQ4_K_M
✓ Measured
Olmo-3.1-32B-Think22.13 tok/s
132 W59°CQ4_K_M
✓ Measured
Qwen3-32B22.04 tok/s
133 W61°CQ4_K_M
✓ Measured
Qwen2.5-32B21.98 tok/s
132 W60°CQ4_K_M
✓ Measured
Qwen2.5-Coder 32B21.92 tok/s
132 W61°CQ4_K_M
✓ Measured
QwQ 32B21.87 tok/s
132 W62°CQ4_K_M
✓ Measured
DeepSeek-R1 Distill 32B21.85 tok/s
135 W63°CQ4_K_M
✓ Measured
Dolphin 2.9.1 Yi 1.5 34B21.12 tok/s
133 W63°CQ4_K_M
✓ Measured
Qwen2.5-Coder 32B (Q3_K_M)18.78 tok/s
133 W61°CQ3_K_M
✓ Measured
Llama 3.3 70B✕ Won't fit needs ~46 GBVRAM-gated at this precision✓ Measured
DeepSeek-R1-Distill-Llama-70B✕ Won't fit needs ~54 GBVRAM-gated at this precision✓ Measured
Hermes-4-70B✕ Won't fit needs ~54 GBVRAM-gated at this precision✓ Measured
Qwen3-Coder-Next-abliterated✕ Won't fit needs ~61 GBVRAM-gated at this precision✓ Measured
Laguna-XS-2.1✕ Won't fit needs ~26 GBVRAM-gated at this precision✓ Measured
Llama-3.3-70B-Instruct-abliterated✕ Won't fit needs ~54 GBVRAM-gated at this precision✓ Measured
Meta-Llama-3.1-70B✕ Won't fit needs ~54 GBVRAM-gated at this precision✓ Measured
Nanbeige4.2-3B✕ Won't fit needs ~4 GBVRAM-gated at this precision✓ Measured
Nemotron-3-Nano-30B-A3B✕ Won't fit needs ~31 GBVRAM-gated at this precision✓ Measured
Qwen3-Coder-Next✕ Won't fit needs ~61 GBVRAM-gated at this precision✓ Measured
Qwen3-Next-80B-A3B-Thinking✕ Won't fit needs ~61 GBVRAM-gated at this precision✓ Measured
Qwen2.5-72B✕ Won't fit needs ~60 GBVRAM-gated at this precision✓ Measured
Qwen3-Coder-Next✕ Won't fit needs ~61 GBVRAM-gated at this precision✓ Measured
Qwen3-Next-80B-A3B-Thinking✕ Won't fit needs ~61 GBVRAM-gated at this precision✓ Measured
Qwen3-Next-80B-A3B✕ Won't fit needs ~61 GBVRAM-gated at this precision✓ Measured

Image Generation images/min 10

Sana 1.6B14.13
PixArt-Sigma XL9.93
Stable Diffusion XL6.38
Z-Image Turbo (1024px)6.32
Stable Diffusion 3.5 Medium4.11
Z-Image Turbo3.45
AuraFlow v0.31.39
WorkloadResultTelemetryData
Sana 1.6B14.13 images/min
149 W48°C
✓ Measured
PixArt-Sigma XL9.93 images/min
149 W55°C
✓ Measured
Stable Diffusion XL6.38 images/min
15.8 GB peak149 W58°C9.4 s/img
✓ Measured
Z-Image Turbo (1024px)6.32 images/min
149 W59°C
✓ Measured
Stable Diffusion 3.5 Medium4.11 images/min
149 W66°C
✓ Measured
Z-Image Turbo3.45 images/min
21.9 GB peak148 W63°C17.4 s/img
✓ Measured
AuraFlow v0.31.39 images/min
150 W68°C
✓ Measured
FLUX.1 dev✕ Won't fit needs ~26 GBVRAM-gated at this precision✓ Measured
FLUX.1 Schnell✕ Won't fit needs ~26 GBVRAM-gated at this precision✓ Measured
Stable Diffusion 3.5 Large✕ Won't fit needs ~24 GBVRAM-gated at this precision✓ Measured

Fine-Tuning train tok/s 4

TinyLlama 1.1B LoRA5499.6
Qwen2.5 1.5B LoRA4214.4
SmolLM2 1.7B LoRA3668.3
Qwen2.5 7B LoRA1239
WorkloadResultTelemetryData
TinyLlama 1.1B LoRA5499.6 train tok/s
147 W66°C
✓ Measured
Qwen2.5 1.5B LoRA4214.4 train tok/s
150 W67°C
✓ Measured
SmolLM2 1.7B LoRA3668.3 train tok/s
150 W67°C
✓ Measured
Qwen2.5 7B LoRA1239 train tok/s
149 W68°C
✓ Measured

LLM Serving serve tok/s 4

TinyLlama 1.1B served3558.2
Qwen2.5 1.5B served2702.5
SmolLM2 1.7B served2242.7
Qwen2.5 7B served858.7
WorkloadResultTelemetryData
TinyLlama 1.1B served3558.2 serve tok/s
154 W34°C
✓ Measured
Qwen2.5 1.5B served2702.5 serve tok/s
160 W37°C
✓ Measured
SmolLM2 1.7B served2242.7 serve tok/s
171 W39°C
✓ Measured
Qwen2.5 7B served858.7 serve tok/s
203 W46°C
✓ Measured

Image to Video clips/min 3

WorkloadResultTelemetryData
CogVideoX-5B I2V✕ Won't fit needs ~24 GBVRAM-gated at this precision✓ Measured
Stable Video Diffusion XT✕ Won't fit needs ~16 GBVRAM-gated at this precision✓ Measured
Wan 2.2 TI2V-5B (image to video)✕ Won't fit needs ~28 GBVRAM-gated at this precision✓ Measured

Image Editing images/min 2

WorkloadResultTelemetryData
FLUX.1 Kontext dev✕ Won't fit needs ~26 GBVRAM-gated at this precision✓ Measured
Qwen-Image-Edit✕ Won't fit needs ~42 GBVRAM-gated at this precision✓ Measured

Depth Estimation images/min 2

WorkloadResultTelemetryData
Depth Anything V2 Small941.18 images/min
55 W33°C
✓ Measured
Depth Anything V2 Large740.74 images/min
55 W33°C
✓ Measured

Segmentation images/min 2

WorkloadResultTelemetryData
SAM ViT-Base398.72 images/min
65 W36°C
✓ Measured
SAM ViT-Huge84.22 images/min
142 W40°C
✓ Measured

Vision Language images/min 2

WorkloadResultTelemetryData
Florence-2 Base210.79 images/min
63 W36°C
✓ Measured
Florence-2 Large117.94 images/min
86 W37°C
✓ Measured

Video Generation frames/s 1

WorkloadResultTelemetryData
Wan 2.2 5B (720p)0.18 frames/s
18.4 GB peak133 W70°C266.2 s/clip
✓ Measured

Image to 3D assets/hour 1

WorkloadResultTelemetryData
TRELLIS Image-to-3D293.6 assets/hour
197 W51°C
✓ Measured

Background Removal images/min 1

WorkloadResultTelemetryData
BiRefNet432.33 images/min
68 W32°C
✓ Measured

Upscaling images/min 1

WorkloadResultTelemetryData
Swin2SR 4x Upscaler22.31 images/min
148 W47°C
✓ Measured

Text to Speech x realtime 1

WorkloadResultTelemetryData
Kokoro TTS 82M101.11 x realtime
66 W36°C
✓ Measured

Music Generation x realtime 1

WorkloadResultTelemetryData
MusicGen Small1.2 x realtime
104 W48°C
✓ Measured

Speech to Text x realtime 1

WorkloadResultTelemetryData
Whisper large-v386.07 x realtime
116 W41°C
✓ Measured

Speculative Decoding x vs solo 1

WorkloadResultTelemetryData
Qwen2.5 1.5B + 0.5B draft0.92 x vs solo✓ Measured
How we measured this. Every result comes from our own pinned, reproducible AI suite, 12 workloads: the Qwen3-4B to Llama-70B LLM ladder (llama.cpp, Q4_K_M), SDXL / Z-Image / FLUX-dev generation, FLUX-Kontext / Qwen-Edit editing, and LTX / Wan video, run first-party on rented hardware with under 0.5% run-to-run variance. Peak VRAM, power draw, temperature and tokens-per-watt are captured per workload. “Won’t fit” rows are real data: where a model exceeds the card’s VRAM at the tested precision we record a hard gate rather than silently dropping to a smaller quant. Measured 2026-07-10 · harness 2.0.0.

NVIDIA A10G specifications

ArchitectureAmpere
CUDA cores9,216
VRAM24GB GDDR6
Memory bus384-bit
Memory bandwidth600 GB/s
Boost clock1,710 MHz
TDP150 W
ProcessSamsung 8nm
InterfacePCIe 4.0 x16
Release date2021-11-01
Launch MSRP$2,800

Verdict, capable, but 24GB sets the ceiling

NVIDIA A10G scores 6.4/100, #39 of 102. It ran 6 of 12; 5 exceeded its 24GB. Every figure here is our own measurement.

Relative performance: where the NVIDIA A10G lands

100% = this card, AI & Machine Learning headline metric (AI Score). #19 of 21 desktop cards in this vertical.

GPURelative%AI Score
NVIDIA L40
305%19.5
NVIDIA A40
277%17.7
NVIDIA A100 40GB SXM4
266%17
NVIDIA A100 40GB PCIe
261%16.7
NVIDIA A10G
100%6.4
NVIDIA L4
78%5
NVIDIA T4
45%2.9

← All AI & Machine Learning GPU rankings

What this card can build

Whole-job timings, composed from our measured per-model results on this card.

WorkflowTimeEnergyBasis
50-image depth pass7 s0.06 Whmeasured
30-minute podcast pass77 s2.03 Whall 3 stages measured
500-image masking run6 min14.07 Whmeasured
20-asset 3D game kit7.6 min21.16 Whall 2 stages measured
24-frame storyboard9 min19.54 Whall 2 stages measured
200-product catalogue cutout9.6 min22.58 Whall 2 stages measured
Full codebase review21.3 min43.41 Whmeasured
10 short social clips50.7 min108.73 Whall 3 stages measured

Can't run: 40-product photo shoot (needs FLUX.1 Kontext dev), 6-panel comic page (needs FLUX.1 dev), 20 long-form articles (needs Llama 3.3 70B), Character sheet, 12 poses (needs FLUX.1 dev), 100-photo restoration batch (needs FLUX.1 Kontext dev), 100-photo restore and enlarge (needs FLUX.1 Kontext dev), 40-product shoot, start to finish (needs FLUX.1 Kontext dev).