NVIDIA A100 40GB PCIe, AI & Machine Learning Benchmarks & Specs

40GB · AI Score 16.7/100 · anchored estimate vs 51 measured cards

16.7 AI Score Includes estimates

We have not run NVIDIA A100 40GB PCIe on our bench. These figures are anchored estimates, interpolated per workload against the 51 GPUs we did measure (confidence: high (sibling silicon)). On Llama 3.1 8B (Q4_K_M) NVIDIA A100 40GB PCIe should deliver about 130 tokens/sec. Stepping up to Qwen3 32B it should hold roughly 35.2 tok/s. Llama 3.3 70B does not fit. It needs roughly 42GB and this card has 40GB. For image generation, SDXL should run near 8.4 it/s, and FLUX.1-dev at 1.85 it/s. 2 of the 12 workloads won't fit on 40GB at the tested precision, Llama 3.3 70B, Qwen-Image-Edit. We publish those as hard gates rather than quietly dropping to a smaller quant.

AI & Machine Learning benchmark results

Text Generation tok/s 5

Qwen3 4B159.3
Llama 3.1 8B130
Qwen2.5-Coder 14B71
Qwen3 32B35.2
WorkloadResultTelemetryData
Qwen3 4B159.3 tok/sestimatedEst.
Llama 3.1 8B130 tok/sestimatedEst.
Qwen2.5-Coder 14B71 tok/sestimatedEst.
Qwen3 32B35.2 tok/sestimatedEst.
Llama 3.3 70B✕ Won't fit VRAM-gated at this precisionEst.

Image Generation images/min 3

Stable Diffusion XL16.8
Z-Image Turbo9.3
FLUX.1 dev3.964
WorkloadResultTelemetryData
Stable Diffusion XL16.8 images/minestimatedEst.
Z-Image Turbo9.3 images/minestimatedEst.
FLUX.1 dev3.96 images/minestimatedEst.

Image Editing images/min 2

WorkloadResultTelemetryData
FLUX.1 Kontext dev1.84 images/minestimatedEst.
Qwen-Image-Edit✕ Won't fit VRAM-gated at this precisionEst.

Video Generation frames/s 2

WorkloadResultTelemetryData
LTX-Video (distilled)8.51 frames/sestimatedEst.
Wan 2.2 5B (720p)0.61 frames/sestimatedEst.
How this estimate is derived. This card hasn’t been through our bench yet, so its numbers are anchored estimates, interpolated from the 51 first-party measured cards (Sibling-anchored to measured A100 80GB: LLM ×0.804 (bandwidth ratio), diffusion ×1.0 (identical compute); VRAM gates at 40GB). The VRAM “won’t fit” gates are exact, since they’re pure capacity limits. Confidence: high (sibling silicon). Estimates are replaced with measured data as more silicon goes through the bench. Full methodology →

NVIDIA A100 40GB PCIe specifications

ArchitectureAmpere
CUDA cores6,912
VRAM40GB HBM2
Memory bus5120-bit
Memory bandwidth1555 GB/s
Boost clock1,410 MHz
TDP250 W
ProcessTSMC 7nm
InterfacePCIe 4.0 x16
Release date2020-06-22
Launch MSRP$10,000

Verdict, capable, but 40GB sets the ceiling

NVIDIA A100 40GB PCIe scores 16.7/100, #24 of 102. It ran 10 of 12; 2 exceeded its 40GB. Figures are anchored estimates, not measurements, we flag that on every row.

Relative performance: where the NVIDIA A100 40GB PCIe lands

100% = this card, AI & Machine Learning headline metric (AI Score). #18 of 21 desktop cards in this vertical.

GPURelative%AI Score
NVIDIA L40S
166%27.7
NVIDIA L40
117%19.5
NVIDIA A40
106%17.7
NVIDIA A100 40GB SXM4
102%17
NVIDIA A100 40GB PCIe
100%16.7
NVIDIA A10G
38%6.4
NVIDIA L4
30%5
NVIDIA T4
17%2.9

← All AI & Machine Learning GPU rankings

What this card can build

Whole-job timings, composed from our measured per-model results on this card.

WorkflowTimeEnergyBasis
24-frame storyboard3.2 minn/aestimate, 0 of 2 stages measured
60-second AI short film4.2 minn/aestimate, 0 of 3 stages measured
6-panel comic page5.1 minn/aestimate, 0 of 3 stages measured
Character sheet, 12 poses6.8 minn/aestimate, 0 of 2 stages measured
Full codebase review14.1 minn/aestimate, 0 of 1 stage measured
10 short social clips14.7 minn/aestimate, 0 of 3 stages measured
40-product photo shoot24.1 minn/aestimate, 0 of 2 stages measured
100-photo restoration batch54.3 minn/aestimate, 0 of 1 stage measured

Can't run: 20 long-form articles (needs Llama 3.3 70B).