Stable Diffusion XL · 53 cards measured first-party · Updated October 2026

How Fast Is Stable Diffusion XL on Each GPU?

SDXL is the image model everyone actually runs. It needs ~8GB minimum and ~12GB to be comfortable, which puts it within reach of most of the market: and unlike our LLM ladder, it rewards raw tensor compute rather than memory bandwidth. That flips the ranking completely.

Benchmarked weights: stabilityai/stable-diffusion-xl-base-1.0

Fastest we measured
NVIDIA B200

NVIDIA B200

46.12 images/min on Stable Diffusion XL, the ceiling. Measured on our bench. 192GB of VRAM, $40,000 at launch.

Pros
  • 46.12 images/min on Stable Diffusion XL
  • 192GB, clears the Stable Diffusion XL floor
  • Rentable by the hour rather than bought
Cons
  • 1000W board rating
  • Datacenter or workstation hardware, not a retail purchase

Best for: Stable Diffusion XL work where you want the ceiling gone rather than the cheapest entry.

Best consumer card
NVIDIA GeForce RTX 5090

NVIDIA GeForce RTX 5090

21.08 images/min on Stable Diffusion XL, fastest card you can buy at retail. Measured on our bench. 32GB of VRAM, $1,999 at launch.

Pros
  • 21.08 images/min on Stable Diffusion XL
  • 32GB, clears the Stable Diffusion XL floor
Cons
  • 575W board rating
Cheapest card that runs it
NVIDIA GeForce RTX 5060

NVIDIA GeForce RTX 5060

4.28 images/min on Stable Diffusion XL, lowest launch price that still fits. Anchored estimate. 8GB of VRAM, $249 at launch.

Pros
  • 4.28 images/min on Stable Diffusion XL
  • 8GB, clears the Stable Diffusion XL floor
Cons
  • 145W board rating
Best value
GeForce RTX 5070

GeForce RTX 5070

7.43 images/min on Stable Diffusion XL, most speed per dollar. Measured on our bench. 12GB of VRAM, $549 at launch. That is 13.53 images/min per $1,000 of launch price.

Pros
  • 7.43 images/min on Stable Diffusion XL
  • 12GB, clears the Stable Diffusion XL floor
Cons
  • 250W board rating
46.12images/min
Fastest: NVIDIA B200
measured
78
Cards that run Stable Diffusion XL
of 83 we have data for
5
Cards that can't run it at all
published as hard gates, not omissions
4251%
Fastest vs slowest that fits
46.12 vs 1.06 images/min

Diffusion is tensor-compute bound. Every step is dense matrix maths, so the ranking tracks tensor throughput and architecture generation, not bandwidth. This is why an Ampere card with plenty of VRAM gets buried by a newer card with less, and it's the opposite of what the LLM charts show. One card, two completely different orderings, depending on the job.

Stable Diffusion XL: speed on every GPU we have data for

NVIDIA B200
46.12 images/min
NVIDIA GH200 Grace Hopper
37.16 images/min
NVIDIA H200
37.16 images/min
NVIDIA B100
36.9 images/min
NVIDIA H100 NVL
34.58 images/min
NVIDIA H100 80GB HBM3
34.58 images/min
NVIDIA H800 80GB
34.58 images/min
NVIDIA H100 PCIe
29.86 images/min
NVIDIA B300
29.2 images/min
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
27.94 images/min
NVIDIA RTX PRO 6000 Blackwell Server Edition
26.32 images/min
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
21.8 images/min
NVIDIA GeForce RTX 5090
21.08 images/min
NVIDIA A100 40GB SXM4
18.58 images/min
NVIDIA RTX PRO 5000 Blackwell
17.46 images/min

Top 15 shown; 63 more cards in the full table below.

Single stream, batch size 1. 37 of the 59 cards on this page were measured first-party by us; the rest are anchored estimates against those measurements and are labelled in the table below.

Efficiency: images/min per 100W drawn

NVIDIA L4
7.19 images/min / 100W
NVIDIA RTX PRO 4500 Blackwell
6.57 images/min / 100W
NVIDIA B200
6.42 images/min / 100W
NVIDIA H100 80GB HBM3
6.42 images/min / 100W
NVIDIA H200
6.15 images/min / 100W
NVIDIA A100 80GB PCIe
6.1 images/min / 100W
NVIDIA B300
5.98 images/min / 100W
NVIDIA RTX PRO 4000 Blackwell
5.9 images/min / 100W
NVIDIA RTX PRO 5000 Blackwell
5.87 images/min / 100W
NVIDIA RTX 5000 Ada Generation
5.41 images/min / 100W
NVIDIA RTX 4000 (Ada Generation)
5.36 images/min / 100W
NVIDIA RTX 2000 Ada Generation
5.23 images/min / 100W
NVIDIA L40S
5.22 images/min / 100W
NVIDIA RTX PRO 6000 Blackwell Server Edition
5.17 images/min / 100W
NVIDIA A100 40GB SXM4
4.95 images/min / 100W

Top 15 shown; 36 more cards in the full table below.

Power is the average pulled during the run, sampled at 1Hz. The fastest card is often not the one here, and for anything left running this is the number that shows up on the bill.

Value: images/min per $1,000 of MSRP

GeForce RTX 5070
13.53 images/min / $1k
NVIDIA GeForce RTX 4070 Super
13.12 images/min / $1k
GeForce RTX 5070 Ti
12.52 images/min / $1k
GeForce RTX 5080
12.25 images/min / $1k
GeForce RTX 5060 Ti
12.17 images/min / $1k
GeForce RTX 4080 Super
11.55 images/min / $1k
NVIDIA GeForce RTX 4070 Ti Super
11.14 images/min / $1k
NVIDIA GeForce RTX 4070
11.02 images/min / $1k
NVIDIA GeForce RTX 4070 Ti
10.83 images/min / $1k
NVIDIA GeForce RTX 5090
10.55 images/min / $1k
NVIDIA GeForce RTX 4090
10.18 images/min / $1k
GeForce RTX 4060
9.3 images/min / $1k
NVIDIA GeForce RTX 4060 Ti 16GB
9.18 images/min / $1k
NVIDIA GeForce RTX 3060
8.94 images/min / $1k
NVIDIA GeForce RTX 4080
8.47 images/min / $1k

Top 15 shown; 37 more cards in the full table below.

Launch price, not street price, so it ages. A speed leaderboard always crowns the most expensive card; this is the counterweight.

Won't fit, Stable Diffusion XL gates these cards outright

NVIDIA GeForce GTX 1660 Super6GB
NVIDIA GeForce GTX 1660 Ti6GB
NVIDIA GeForce GTX 16606GB
NVIDIA GeForce RTX 20606GB
NVIDIA RTX A20006GB
GPUVRAMWhy it fails
NVIDIA GeForce GTX 1660 Super6GBNeeds ~11GB VRAM
NVIDIA GeForce GTX 1660 Ti6GBNeeds ~11GB VRAM
NVIDIA GeForce GTX 16606GBNeeds ~11GB VRAM
NVIDIA GeForce RTX 20606GBNeeds ~11GB VRAM
NVIDIA RTX A20006GBNeeds ~11GB VRAM

No driver update fixes a VRAM ceiling.

Full Stable Diffusion XL leaderboard, every card that runs it

NVIDIA B20046.12 images/min
NVIDIA GH200 Grace Hopper37.16 images/min
NVIDIA H20037.16 images/min
NVIDIA B10036.9 images/min
NVIDIA H100 NVL34.58 images/min
NVIDIA H100 80GB HBM334.58 images/min
NVIDIA H800 80GB34.58 images/min
NVIDIA H100 PCIe29.86 images/min
NVIDIA B30029.2 images/min
NVIDIA RTX PRO 6000 Blackwell Workstation Edition27.94 images/min
NVIDIA RTX PRO 6000 Blackwell Server Edition26.32 images/min
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition21.8 images/min
NVIDIA GeForce RTX 509021.08 images/min
NVIDIA A100 40GB SXM418.58 images/min
NVIDIA RTX PRO 5000 Blackwell17.46 images/min
NVIDIA L40S17.02 images/min
NVIDIA A100 40GB PCIe16.8 images/min
NVIDIA A100 80GB PCIe16.8 images/min
NVIDIA A100 80GB SXM416.36 images/min
NVIDIA A800 80GB16.36 images/min
NVIDIA GeForce RTX 409016.28 images/min
NVIDIA RTX PRO 4500 Blackwell13.1 images/min
NVIDIA RTX 5000 Ada Generation12.74 images/min
GeForce RTX 508012.24 images/min
GeForce RTX 4080 Super11.54 images/min
NVIDIA L4011.2 images/min
NVIDIA RTX 6000 Ada Generation10.7 images/min
NVIDIA GeForce RTX 408010.16 images/min
NVIDIA RTX A600010.02 images/min
NVIDIA RTX 5880 Ada Generation9.96 images/min
GeForce RTX 5070 Ti9.38 images/min
NVIDIA A409.2 images/min
NVIDIA GeForce RTX 4070 Ti Super8.9 images/min
NVIDIA GeForce RTX 4070 Ti8.65 images/min
NVIDIA RTX PRO 4000 Blackwell8.56 images/min
NVIDIA RTX A55008.51 images/min
NVIDIA GeForce RTX 3090 Ti8.42 images/min
NVIDIA RTX 4500 Ada Generation8.32 images/min
NVIDIA GeForce RTX 3080 Ti8.03 images/min
NVIDIA GeForce RTX 4070 Super7.86 images/min
NVIDIA GeForce RTX 30907.46 images/min
NVIDIA RTX A50007.46 images/min
GeForce RTX 50707.43 images/min
NVIDIA GeForce RTX 40706.6 images/min
NVIDIA RTX 4000 (Ada Generation)6.56 images/min
NVIDIA Titan RTX6.56 images/min
NVIDIA RTX A45006.5 images/min
NVIDIA A10G6.38 images/min
NVIDIA Quadro RTX 80006.18 images/min
NVIDIA Quadro RTX 6000 (Turing)5.94 images/min
NVIDIA RTX A40005.24 images/min
GeForce RTX 5060 Ti5.22 images/min
NVIDIA L45.18 images/min
NVIDIA GeForce RTX 30804.76 images/min
NVIDIA GeForce RTX 4060 Ti 16GB4.58 images/min
NVIDIA TITAN V4.44 images/min
NVIDIA GeForce RTX 50604.28 images/min
NVIDIA GeForce RTX 2080 Ti Founders Edition4.18 images/min
NVIDIA GeForce RTX 3070 Ti3.76 images/min
NVIDIA RTX 2000 Ada Generation3.58 images/min
NVIDIA GeForce RTX 2080 Super3.28 images/min
NVIDIA Quadro RTX 50003.14 images/min
NVIDIA GeForce RTX 3070 Founders Edition3.04 images/min
NVIDIA GeForce RTX 30602.94 images/min
NVIDIA GeForce RTX 50502.91 images/min
GeForce RTX 40602.78 images/min
NVIDIA GeForce RTX 2080 Founders Edition2.74 images/min
NVIDIA GeForce RTX 2070 SUPER2.6 images/min
NVIDIA T42.36 images/min
NVIDIA GeForce RTX 20702.2 images/min
NVIDIA GeForce RTX 3060 Ti2.2 images/min
NVIDIA GeForce RTX 2060 Super2.11 images/min
NVIDIA GeForce RTX 30502.0 images/min
NVIDIA TITAN Xp1.8 images/min
GeForce GTX 1080 Ti1.68 images/min
NVIDIA TITAN X (Pascal)1.65 images/min
NVIDIA GeForce GTX 10801.08 images/min
NVIDIA GeForce GTX 1070 Ti1.06 images/min
GPUResultVRAMSource
NVIDIA B20046.12 images/min192GBMeasured
NVIDIA GH200 Grace Hopper37.16 images/min141GBEstimated
NVIDIA H20037.16 images/min141GBMeasured
NVIDIA B10036.9 images/min192GBEstimated
NVIDIA H100 NVL34.58 images/min94GBEstimated
NVIDIA H100 80GB HBM334.58 images/min80GBMeasured
NVIDIA H800 80GB34.58 images/min80GBEstimated
NVIDIA H100 PCIe29.86 images/min80GBEstimated
NVIDIA B30029.2 images/min288GBMeasured
NVIDIA RTX PRO 6000 Blackwell Workstation Edition27.94 images/min96GBMeasured
NVIDIA RTX PRO 6000 Blackwell Server Edition26.32 images/min96GBMeasured
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition21.8 images/min96GBEstimated
NVIDIA GeForce RTX 509021.08 images/min32GBMeasured
NVIDIA A100 40GB SXM418.58 images/min40GBMeasured
NVIDIA RTX PRO 5000 Blackwell17.46 images/min48GBMeasured
NVIDIA L40S17.02 images/min48GBMeasured
NVIDIA A100 40GB PCIe16.8 images/min40GBEstimated
NVIDIA A100 80GB PCIe16.8 images/min80GBMeasured
NVIDIA A100 80GB SXM416.36 images/min80GBMeasured
NVIDIA A800 80GB16.36 images/min80GBEstimated
NVIDIA GeForce RTX 409016.28 images/min24GBMeasured
NVIDIA RTX PRO 4500 Blackwell13.1 images/min32GBMeasured
NVIDIA RTX 5000 Ada Generation12.74 images/min32GBMeasured
GeForce RTX 508012.24 images/min16GBMeasured
GeForce RTX 4080 Super11.54 images/min16GBMeasured
NVIDIA L4011.2 images/min48GBMeasured
NVIDIA RTX 6000 Ada Generation10.7 images/min48GBMeasured
NVIDIA GeForce RTX 408010.16 images/min16GBMeasured
NVIDIA RTX A600010.02 images/min48GBMeasured
NVIDIA RTX 5880 Ada Generation9.96 images/min48GBEstimated
GeForce RTX 5070 Ti9.38 images/min16GBMeasured
NVIDIA A409.2 images/min48GBMeasured
NVIDIA GeForce RTX 4070 Ti Super8.9 images/min16GBMeasured
NVIDIA GeForce RTX 4070 Ti8.65 images/min12GBMeasured
NVIDIA RTX PRO 4000 Blackwell8.56 images/min24GBMeasured
NVIDIA RTX A55008.51 images/min24GBEstimated
NVIDIA GeForce RTX 3090 Ti8.42 images/min24GBMeasured
NVIDIA RTX 4500 Ada Generation8.32 images/min24GBEstimated
NVIDIA GeForce RTX 3080 Ti8.03 images/min12GBMeasured
NVIDIA GeForce RTX 4070 Super7.86 images/min12GBMeasured
NVIDIA GeForce RTX 30907.46 images/min24GBMeasured
NVIDIA RTX A50007.46 images/min24GBMeasured
GeForce RTX 50707.43 images/min12GBMeasured
NVIDIA GeForce RTX 40706.6 images/min12GBMeasured
NVIDIA RTX 4000 (Ada Generation)6.56 images/min20GBMeasured
NVIDIA Titan RTX6.56 images/min24GBMeasured
NVIDIA RTX A45006.5 images/min20GBMeasured
NVIDIA A10G6.38 images/min24GBMeasured
NVIDIA Quadro RTX 80006.18 images/min48GBMeasured
NVIDIA Quadro RTX 6000 (Turing)5.94 images/min24GBMeasured
NVIDIA RTX A40005.24 images/min16GBMeasured
GeForce RTX 5060 Ti5.22 images/min16GBMeasured
NVIDIA L45.18 images/min24GBMeasured
NVIDIA GeForce RTX 30804.76 images/min10GBMeasured
NVIDIA GeForce RTX 4060 Ti 16GB4.58 images/min16GBMeasured
NVIDIA TITAN V4.44 images/min12GBEstimated
NVIDIA GeForce RTX 50604.28 images/min8GBEstimated
NVIDIA GeForce RTX 2080 Ti Founders Edition4.18 images/min11GBMeasured
NVIDIA GeForce RTX 3070 Ti3.76 images/min8GBMeasured
NVIDIA RTX 2000 Ada Generation3.58 images/min16GBMeasured
NVIDIA GeForce RTX 2080 Super3.28 images/min8GBEstimated
NVIDIA Quadro RTX 50003.14 images/min16GBEstimated
NVIDIA GeForce RTX 3070 Founders Edition3.04 images/min8GBMeasured
NVIDIA GeForce RTX 30602.94 images/min12GBMeasured
NVIDIA GeForce RTX 50502.91 images/min8GBEstimated
GeForce RTX 40602.78 images/min8GBMeasured
NVIDIA GeForce RTX 2080 Founders Edition2.74 images/min8GBEstimated
NVIDIA GeForce RTX 2070 SUPER2.6 images/min8GBEstimated
NVIDIA T42.36 images/min16GBMeasured
NVIDIA GeForce RTX 20702.2 images/min8GBEstimated
NVIDIA GeForce RTX 3060 Ti2.2 images/min8GBMeasured
NVIDIA GeForce RTX 2060 Super2.11 images/min8GBEstimated
NVIDIA GeForce RTX 30502.0 images/min8GBEstimated
NVIDIA TITAN Xp1.8 images/min12GBEstimated
GeForce GTX 1080 Ti1.68 images/min11GBEstimated
NVIDIA TITAN X (Pascal)1.65 images/min12GBEstimated
NVIDIA GeForce GTX 10801.08 images/min8GBEstimated
NVIDIA GeForce GTX 1070 Ti1.06 images/min8GBEstimated

Tap any column to sort. Measured = we rented and ran this card ourselves. Estimated = interpolated against our measured anchors, never blended silently.

Because this workload is tensor-compute bound, the ranking tracks architecture generation and tensor throughput rather than memory bandwidth, the reverse of our LLM charts. The same two cards can swap places entirely depending on which of these pages you're reading. That's the reason we run twelve workloads instead of publishing one score. A GPU isn't fast or slow. It's fast at some things and gated out of others, and which of those matters depends entirely on what you're actually going to run.

Our verdict

NVIDIA B200 tops our Stable Diffusion XL leaderboard at 46.12 images/min (measured), 4251% of the way clear of the slowest card that still fits. But the number that decides most purchases isn't on the chart. It's the 5 cards that can't run Stable Diffusion XL at all. This is a compute workload: buy architecture generation, not raw VRAM, as long as you clear the floor first.

FAQ

What is the fastest GPU for Stable Diffusion XL?
NVIDIA B200, at 46.12 images/min on our bench, a first-party measurement. It carries 192GB of VRAM. Of the 83 cards we have Stable Diffusion XL data for, 78 can run it at all.
How much VRAM do I need for Stable Diffusion XL?
~8GB minimum, ~12GB for full-fat. One of the few workloads in our suite that nearly everything can run.
Why does the Stable Diffusion XL ranking look different from your other benchmarks?
Because this workload is tensor-compute bound, the ranking tracks architecture generation and tensor throughput rather than memory bandwidth, the reverse of our LLM charts. The same two cards can swap places entirely depending on which of these pages you're reading. That's why we publish twelve separate workloads rather than one blended score, the ordering genuinely changes depending on the job.
Are these Stable Diffusion XL numbers measured or estimated?
Both, and every row says which. 53 of the 83 cards here were rented and run by us on the same harness. The remainder are anchored estimates interpolated per workload against those measurements. We never blend the two silently, if a row says Estimated, we have not run that card.
Can I rent a GPU to run Stable Diffusion XL instead of buying one?
Yes, and for the cards at the top of this leaderboard it's the only realistic option, most of them have no retail channel at all. It's also how we got these numbers: we rented the hardware by the hour rather than buying it. That's worth considering before you spend on a card to find out whether it's fast enough.
Why publish cards that can't run Stable Diffusion XL?
Because it's the most useful thing we know. A card that can't load a model doesn't run it slowly, it doesn't run it. Most benchmark sites leave that as a blank cell or quietly drop to a smaller quantisation to produce a number. We publish it as a hard gate and score it zero, because 'this card cannot do the thing you want' is the answer to the question you were actually asking.

How we test

Every ranking on this page comes from our own benchmark runs, not vendor claims. Cards marked Measured were rented and run by us; cards marked Estimated are interpolated per workload against those measured anchors and are labelled on every row, we never blend the two silently. LLMs run on llama.cpp (llama-bench) at Q4_K_M with -p 512 -n 128. Diffusion and video run on diffusers/ComfyUI at BF16, with SDXL at FP16. Each workload gets a warmup pass plus multiple timed runs (5 for small LLMs, 3 for large models and images, 2 for video); we publish the mean as the result and the minimum as the 1% low. Run-to-run variance is under 0.5%. Telemetry, power, temperature, utilisation, clocks, peak VRAM, is sampled at 1 Hz for the duration of every run. Where a model exceeds a card's VRAM we publish a hard won't-fit result rather than quietly dropping to a smaller quantisation. A card that can't run a model scores zero on it. Silently swapping precision to make a number appear would make every number on this site meaningless. All figures are single-GPU, single-stream, batch-size-1. That is the honest way to measure what one card does for one user, and it is deliberately not how a datacenter serves a model. Vendor and MLPerf figures use large batches across many GPUs and will be far higher. Neither is wrong, they answer different questions. Ours answers 'what will this card do for me'.