NVIDIA RTX A6000, AI & Machine Learning Benchmarks & Specs

48GB · AI Score 21.0/100 · first-party measured on 12 AI workloads

21 AI Score ✓ Measured

Every number on this page is first-party: NVIDIA RTX A6000 was run on our pinned 12-workload AI suite, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) NVIDIA RTX A6000 delivers about 124.87 tokens/sec. Stepping up to Qwen3 32B it holds roughly 32.18 tok/s. The full Llama 3.3 70B still runs, at about 15.73 tok/s. For image generation, SDXL runs at 5.01 it/s, and FLUX.1-dev at 1.01 it/s. All 12 workloads fit in 48GB. There is no model in our suite this card has to turn down. NVIDIA RTX A6000 isn't a retail purchase for most people. It's rented by the hour. You can run this exact card on RunPod.

AI & Machine Learning benchmark results

Text Generation tok/s 18

LFM2.5 2.6B279.08
MiniCPM5 2B270.65
Granite 4.1 3B202.59
Nemotron 3 Nano 4B190.78
Qwen3 30B A3B187.52
Qwen3 4B184.34
DeepSeek Coder 7B Instruct v1.5140.44
Gemma 4 26B A4B138.08
Llama 3.1 8B124.87
Llama 3 8B124.54
Qwen3 8B120.24
Ornith 1.5 9B106.49
WorkloadResultTelemetryData
MiniCPM5 2B270.65 tok/s
159 W50°CQ4_K_M
✓ Measured
LFM2.5 2.6B279.08 tok/s
154 W48°CQ4_K_M
✓ Measured
Granite 4.1 3B202.59 tok/s
181 W55°CQ4_K_M
✓ Measured
Nemotron 3 Nano 4B190.78 tok/s
184 W52°CQ4_K_M
✓ Measured
Qwen3 4B184.34 tok/s
2.7 GB peak144 W48°C1.28 tok/WQ4_K_M
✓ Measured
DeepSeek Coder 7B Instruct v1.5140.44 tok/s
196 W54°CQ4_K_M
✓ Measured
Llama 3 8B124.54 tok/s
195 W55°CQ4_K_M
✓ Measured
Llama 3.1 8B124.87 tok/s
4.7 GB peak169 W52°C0.74 tok/WQ4_K_M
✓ Measured
Qwen3 8B120.24 tok/s
195 W55°CQ4_K_M
✓ Measured
Nemotron Nano 9B v290.62 tok/s
209 W54°CQ4_K_M
✓ Measured
Ornith 1.5 9B106.49 tok/s
198 W54°CQ4_K_M
✓ Measured
Gemma 4 12B75.82 tok/s
206 W55°CQ4_K_M
✓ Measured
Qwen2.5-Coder 14B68.51 tok/s
8.7 GB peak171 W58°C0.4 tok/WQ4_K_M
✓ Measured
Qwen3 14B70.24 tok/s
212 W55°CQ4_K_M
✓ Measured
Gemma 4 26B A4B138.08 tok/s
156 W53°CQ4_K_M
✓ Measured
Qwen3 30B A3B187.52 tok/s
144 W53°CQ4_K_M
✓ Measured
Qwen3 32B32.18 tok/s
18.7 GB peak126 W60°C0.26 tok/WQ4_K_M
✓ Measured
Llama 3.3 70B15.73 tok/s
39.8 GB peak114 W63°C0.14 tok/WQ4_K_M
✓ Measured

Image Generation images/min 4

Stable Diffusion XL10.02
Z-Image Turbo5.4
FLUX.1 dev2.164
Chroma1-HD0.62
WorkloadResultTelemetryData
Stable Diffusion XL10.02 images/min
15.8 GB peak297 W59°C6 s/img
✓ Measured
Z-Image Turbo5.4 images/min
25.6 GB peak296 W67°C11.1 s/img
✓ Measured
Chroma1-HD0.62 images/min
299 W88°C
✓ Measured
FLUX.1 dev2.16 images/min
36.5 GB peak298 W76°C27.7 s/img
✓ Measured

Image Editing images/min 1

WorkloadResultTelemetryData
FLUX.1 Kontext dev1.05 images/min
35.3 GB peak296 W79°C57.7 s/img
✓ Measured
How we measured this. Every result comes from our own pinned, reproducible AI suite, 12 workloads: the Qwen3-4B to Llama-70B LLM ladder (llama.cpp, Q4_K_M), SDXL / Z-Image / FLUX-dev generation, FLUX-Kontext / Qwen-Edit editing, and LTX / Wan video, run first-party on rented hardware with under 0.5% run-to-run variance. Peak VRAM, power draw, temperature and tokens-per-watt are captured per workload. “Won’t fit” rows are real data: where a model exceeds the card’s VRAM at the tested precision we record a hard gate rather than silently dropping to a smaller quant.

NVIDIA RTX A6000 specifications

ArchitectureAmpere
CUDA cores10,752
VRAM48GB GDDR6
Memory bus384-bit
Memory bandwidth768 GB/s
Boost clock1,860 MHz
TDP300 W
Process8nm
InterfacePCIe 4.0 x16
Release date2020-10-05
Launch MSRP$4,650

Verdict, NVIDIA RTX A6000 on real AI workloads

NVIDIA RTX A6000 scores 21.0/100, #20 of 102. It ran all 12 workloads. Every figure here is our own measurement.

Relative performance: where the NVIDIA RTX A6000 lands

100% = this card, AI & Machine Learning headline metric (AI Score). #4 of 20 workstation cards in this vertical.

GPURelative%AI Score
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
253%53.1
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
207%43.5
NVIDIA RTX PRO 5000 Blackwell
150%31.6
NVIDIA RTX A6000
100%21
NVIDIA RTX PRO 4500 Blackwell
61%12.9
NVIDIA RTX A5500
40%8.4
NVIDIA Quadro RTX 8000
39%8.2
NVIDIA RTX PRO 4000 Blackwell
36%7.6

← All AI & Machine Learning GPU rankings

The silicon

Transistors28,300 million
Die size628.4 mm²
Process node8 nm
Fabricated bySamsung
Transistor density45 million per mm²

Denser than 81% of the 746 cards we have silicon data for. Density is the clearest measure of what a process node bought: a card that gained it without growing the die got its speed from the fab rather than the architecture.

Silicon figures from Wikipedia (CC BY-SA 4.0). Benchmarks on this page are our own. Compare every chip.

What this card can build

Whole-job timings, composed from our measured per-model results on this card.

WorkflowTimeEnergyBasis
24-frame storyboard7.1 min23.42 Whall 2 stages measured
6-panel comic page12.1 min42.76 Whall 3 stages measured
Full codebase review14.6 min41.58 Whmeasured
Character sheet, 12 poses15.2 min58.73 Whall 2 stages measured
Long-form article batch29.7 min56.52 Whmeasured
Product photo shoot45.5 min207.87 Whall 2 stages measured
Photo restoration batch1 h 37 min470.32 Whmeasured

Rent or buy?

This card is $4,650 to buy. The cheapest listed rate on RunPod is $0.330/hour, but that is the floor: we budget $0.396/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 11,742 GPU-hours. Below it you are paying for idle silicon.

How you would use itGPU-hours a yearRental cost a yearTime to break even
2 hours a day, hobby730$28916.1 years
8 hours a day, working on it2,920$1,1564.0 years
24/7, always-on agent8,760$3,4691.3 years

At hobby usage this card is very unlikely to pay for itself before it is superseded. Rent it. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.

Rental price

$0.330/hr+0.0% since 2026-08-14low $0.330 · high $0.330

Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.