Stable Diffusion XL · 37 cards measured first-party · Updated July 2026

How Fast Is Stable Diffusion XL on Each GPU?

SDXL is the image model everyone actually runs. It needs ~8GB minimum and ~12GB to be comfortable, which puts it within reach of most of the market: and unlike our LLM ladder, it rewards raw tensor compute rather than memory bandwidth. That flips the ranking completely.

Benchmarked weights: stabilityai/stable-diffusion-xl-base-1.0

Fastest we measured
NVIDIA B200

NVIDIA B200

23.06 it/s on Stable Diffusion XL. Measured on our bench. 192GB of VRAM, 1000W board rating. AI Score 78.0/100 across our full 12-workload suite.

Pros
  • 23.06 it/s on Stable Diffusion XL
  • 192GB, clears the Stable Diffusion XL floor
  • Rentable by the hour rather than bought
Cons
  • 1000W board rating
  • Datacenter or workstation hardware, not a retail purchase

Best for: Stable Diffusion XL work where you want the ceiling gone rather than the cheapest entry.

Runner-up
NVIDIA H200

NVIDIA H200

18.58 it/s on Stable Diffusion XL. Measured on our bench. 141GB of VRAM, 700W board rating. AI Score 65.0/100 across our full 12-workload suite.

Pros
  • 18.58 it/s on Stable Diffusion XL
  • 141GB, clears the Stable Diffusion XL floor
  • Rentable by the hour rather than bought
Cons
  • 700W board rating
  • Datacenter or workstation hardware, not a retail purchase

Best for: Stable Diffusion XL work where you want the ceiling gone rather than the cheapest entry.

Third
NVIDIA H100 NVL

NVIDIA H100 NVL

17.29 it/s on Stable Diffusion XL. Anchored estimate. 94GB of VRAM, 400W board rating. AI Score 67.0/100 across our full 12-workload suite.

Pros
  • 17.29 it/s on Stable Diffusion XL
  • 94GB, clears the Stable Diffusion XL floor
  • Rentable by the hour rather than bought
Cons
  • 400W board rating
  • Datacenter or workstation hardware, not a retail purchase

Best for: Stable Diffusion XL work where you want the ceiling gone rather than the cheapest entry.

23.06it/s
Fastest: NVIDIA B200
measured
58
Cards that run Stable Diffusion XL
of 59 we have data for
1
Cards that can't run it at all
published as hard gates, not omissions
3719%
Fastest vs slowest that fits
23.06 vs 0.62 it/s

Diffusion is tensor-compute bound. Every step is dense matrix maths, so the ranking tracks tensor throughput and architecture generation, not bandwidth. This is why an Ampere card with plenty of VRAM gets buried by a newer card with less, and it's the opposite of what the LLM charts show. One card, two completely different orderings, depending on the job.

Stable Diffusion XL, the 12 fastest cards we have data for

NVIDIA B200
23.06 it/s
NVIDIA GH200 Grace Hopper
18.58 it/s
NVIDIA H200
18.58 it/s
NVIDIA B100
18.45 it/s
NVIDIA H100 NVL
17.29 it/s
NVIDIA H100 80GB HBM3
17.29 it/s
NVIDIA H800 80GB
17.29 it/s
NVIDIA H100 PCIe
14.93 it/s
NVIDIA B300
14.6 it/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
13.97 it/s
NVIDIA RTX PRO 6000 Blackwell Server Edition
13.16 it/s
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
10.9 it/s

Single stream, batch size 1. 37 of the 59 cards on this page were measured first-party by us; the rest are anchored estimates against those measurements and are labelled in the table below.

Won't fit, Stable Diffusion XL gates these cards outright

GPUVRAMWhy it fails
NVIDIA GeForce RTX 20606GBNeeds needs ~8GB VRAM

No driver update fixes a VRAM ceiling.

Full Stable Diffusion XL leaderboard, every card that runs it

NVIDIA B20023.06 it/s
NVIDIA GH200 Grace Hopper18.58 it/s
NVIDIA H20018.58 it/s
NVIDIA B10018.45 it/s
NVIDIA H100 NVL17.29 it/s
NVIDIA H100 80GB HBM317.29 it/s
NVIDIA H800 80GB17.29 it/s
NVIDIA H100 PCIe14.93 it/s
NVIDIA B30014.6 it/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition13.97 it/s
NVIDIA RTX PRO 6000 Blackwell Server Edition13.16 it/s
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition10.9 it/s
NVIDIA GeForce RTX 509010.54 it/s
NVIDIA RTX PRO 5000 Blackwell8.73 it/s
NVIDIA L40S8.51 it/s
NVIDIA A100 40GB PCIe8.4 it/s
NVIDIA A100 80GB PCIe8.4 it/s
NVIDIA A100 40GB SXM48.18 it/s
NVIDIA A100 80GB SXM48.18 it/s
NVIDIA A800 80GB8.18 it/s
NVIDIA GeForce RTX 40908.14 it/s
NVIDIA RTX 5880 Ada Generation6.92 it/s
NVIDIA RTX PRO 4500 Blackwell6.55 it/s
NVIDIA RTX 5000 Ada Generation6.37 it/s
NVIDIA L405.6 it/s
NVIDIA RTX 6000 Ada Generation5.35 it/s
NVIDIA GeForce RTX 40805.08 it/s
NVIDIA RTX A60005.01 it/s
GeForce RTX 5070 Ti4.69 it/s
GeForce RTX 50804.43 it/s
NVIDIA RTX PRO 4000 Blackwell4.28 it/s
NVIDIA RTX 4500 Ada Generation4.16 it/s
NVIDIA GeForce RTX 30903.73 it/s
NVIDIA RTX A50003.73 it/s
NVIDIA RTX 4000 (Ada Generation)3.28 it/s
NVIDIA RTX A45003.25 it/s
NVIDIA A10G3.19 it/s
NVIDIA Quadro RTX 80003.09 it/s
NVIDIA RTX A55003.06 it/s
NVIDIA Quadro RTX 6000 (Turing)2.97 it/s
NVIDIA RTX A40002.62 it/s
NVIDIA L42.59 it/s
AMD Radeon Pro W79002.5 it/s
NVIDIA GeForce RTX 30802.38 it/s
NVIDIA GeForce RTX 4060 Ti2.29 it/s
NVIDIA RTX 2000 Ada Generation1.79 it/s
NVIDIA GeForce RTX 2080 Super1.64 it/s
NVIDIA Quadro RTX 50001.57 it/s
NVIDIA GeForce RTX 3070 Founders Edition1.52 it/s
NVIDIA GeForce RTX 30601.47 it/s
GeForce RTX 40601.39 it/s
NVIDIA GeForce RTX 2080 Founders Edition1.37 it/s
AMD Radeon Pro W68001.3 it/s
AMD Radeon RX 6900 XT1.3 it/s
NVIDIA GeForce RTX 2070 SUPER1.11 it/s
NVIDIA GeForce RTX 3060 Ti1.1 it/s
NVIDIA GeForce RTX 20700.7 it/s
NVIDIA GeForce RTX 2060 Super0.62 it/s
GPUResultVRAMSource
NVIDIA B20023.06 it/s192GBMeasured
NVIDIA GH200 Grace Hopper18.58 it/s141GBEstimated
NVIDIA H20018.58 it/s141GBMeasured
NVIDIA B10018.45 it/s192GBEstimated
NVIDIA H100 NVL17.29 it/s94GBEstimated
NVIDIA H100 80GB HBM317.29 it/s80GBMeasured
NVIDIA H800 80GB17.29 it/s80GBEstimated
NVIDIA H100 PCIe14.93 it/s80GBEstimated
NVIDIA B30014.6 it/s288GBMeasured
NVIDIA RTX PRO 6000 Blackwell Workstation Edition13.97 it/s96GBMeasured
NVIDIA RTX PRO 6000 Blackwell Server Edition13.16 it/s96GBMeasured
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition10.9 it/s96GBEstimated
NVIDIA GeForce RTX 509010.54 it/s32GBMeasured
NVIDIA RTX PRO 5000 Blackwell8.73 it/s48GBMeasured
NVIDIA L40S8.51 it/s48GBMeasured
NVIDIA A100 40GB PCIe8.4 it/s40GBEstimated
NVIDIA A100 80GB PCIe8.4 it/s80GBMeasured
NVIDIA A100 40GB SXM48.18 it/s40GBEstimated
NVIDIA A100 80GB SXM48.18 it/s80GBMeasured
NVIDIA A800 80GB8.18 it/s80GBEstimated
NVIDIA GeForce RTX 40908.14 it/s24GBMeasured
NVIDIA RTX 5880 Ada Generation6.92 it/s48GBEstimated
NVIDIA RTX PRO 4500 Blackwell6.55 it/s32GBMeasured
NVIDIA RTX 5000 Ada Generation6.37 it/s32GBMeasured
NVIDIA L405.6 it/s48GBMeasured
NVIDIA RTX 6000 Ada Generation5.35 it/s48GBMeasured
NVIDIA GeForce RTX 40805.08 it/s16GBMeasured
NVIDIA RTX A60005.01 it/s48GBMeasured
GeForce RTX 5070 Ti4.69 it/s16GBMeasured
GeForce RTX 50804.43 it/s16GBMeasured
NVIDIA RTX PRO 4000 Blackwell4.28 it/s24GBMeasured
NVIDIA RTX 4500 Ada Generation4.16 it/s24GBEstimated
NVIDIA GeForce RTX 30903.73 it/s24GBMeasured
NVIDIA RTX A50003.73 it/s24GBMeasured
NVIDIA RTX 4000 (Ada Generation)3.28 it/s20GBMeasured
NVIDIA RTX A45003.25 it/s20GBMeasured
NVIDIA A10G3.19 it/s24GBMeasured
NVIDIA Quadro RTX 80003.09 it/s48GBMeasured
NVIDIA RTX A55003.06 it/s24GBEstimated
NVIDIA Quadro RTX 6000 (Turing)2.97 it/s24GBMeasured
NVIDIA RTX A40002.62 it/s16GBMeasured
NVIDIA L42.59 it/s24GBMeasured
AMD Radeon Pro W79002.5 it/s48GBEstimated
NVIDIA GeForce RTX 30802.38 it/s10GBMeasured
NVIDIA GeForce RTX 4060 Ti2.29 it/s16GBMeasured
NVIDIA RTX 2000 Ada Generation1.79 it/s16GBMeasured
NVIDIA GeForce RTX 2080 Super1.64 it/s8GBEstimated
NVIDIA Quadro RTX 50001.57 it/s16GBEstimated
NVIDIA GeForce RTX 3070 Founders Edition1.52 it/s8GBMeasured
NVIDIA GeForce RTX 30601.47 it/s12GBMeasured
GeForce RTX 40601.39 it/s8GBMeasured
NVIDIA GeForce RTX 2080 Founders Edition1.37 it/s8GBEstimated
AMD Radeon Pro W68001.3 it/s32GBEstimated
AMD Radeon RX 6900 XT1.3 it/s16GBEstimated
NVIDIA GeForce RTX 2070 SUPER1.11 it/s8GBEstimated
NVIDIA GeForce RTX 3060 Ti1.1 it/s8GBMeasured
NVIDIA GeForce RTX 20700.7 it/s8GBEstimated
NVIDIA GeForce RTX 2060 Super0.62 it/s8GBEstimated

Tap any column to sort. Measured = we rented and ran this card ourselves. Estimated = interpolated against our measured anchors, never blended silently.

Because this workload is tensor-compute bound, the ranking tracks architecture generation and tensor throughput rather than memory bandwidth, the reverse of our LLM charts. The same two cards can swap places entirely depending on which of these pages you're reading. That's the reason we run twelve workloads instead of publishing one score. A GPU isn't fast or slow. It's fast at some things and gated out of others, and which of those matters depends entirely on what you're actually going to run.

Our verdict

NVIDIA B200 tops our Stable Diffusion XL leaderboard at 23.06 it/s (measured), 3719% of the way clear of the slowest card that still fits. But the number that decides most purchases isn't on the chart. It's the 1 cards that can't run Stable Diffusion XL at all. This is a compute workload: buy architecture generation, not raw VRAM, as long as you clear the floor first.

FAQ

What is the fastest GPU for Stable Diffusion XL?
NVIDIA B200, at 23.06 it/s on our bench, a first-party measurement. It carries 192GB of VRAM. Of the 59 cards we have Stable Diffusion XL data for, 58 can run it at all.
How much VRAM do I need for Stable Diffusion XL?
~8GB minimum, ~12GB for full-fat. One of the few workloads in our suite that nearly everything can run.
Why does the Stable Diffusion XL ranking look different from your other benchmarks?
Because this workload is tensor-compute bound, the ranking tracks architecture generation and tensor throughput rather than memory bandwidth, the reverse of our LLM charts. The same two cards can swap places entirely depending on which of these pages you're reading. That's why we publish twelve separate workloads rather than one blended score, the ordering genuinely changes depending on the job.
Are these Stable Diffusion XL numbers measured or estimated?
Both, and every row says which. 37 of the 59 cards here were rented and run by us on the same harness. The remainder are anchored estimates interpolated per workload against those measurements. We never blend the two silently, if a row says Estimated, we have not run that card.
Can I rent a GPU to run Stable Diffusion XL instead of buying one?
Yes, and for the cards at the top of this leaderboard it's the only realistic option, most of them have no retail channel at all. It's also how we got these numbers: we rented the hardware by the hour rather than buying it. That's worth considering before you spend on a card to find out whether it's fast enough.
Why publish cards that can't run Stable Diffusion XL?
Because it's the most useful thing we know. A card that can't load a model doesn't run it slowly, it doesn't run it. Most benchmark sites leave that as a blank cell or quietly drop to a smaller quantisation to produce a number. We publish it as a hard gate and score it zero, because 'this card cannot do the thing you want' is the answer to the question you were actually asking.

How we test

Every ranking on this page comes from our own benchmark runs, not vendor claims. Cards marked Measured were rented and run by us; cards marked Estimated are interpolated per workload against those measured anchors and are labelled on every row, we never blend the two silently. LLMs run on llama.cpp (llama-bench) at Q4_K_M with -p 512 -n 128. Diffusion and video run on diffusers/ComfyUI at BF16, with SDXL at FP16. Each workload gets a warmup pass plus multiple timed runs (5 for small LLMs, 3 for large models and images, 2 for video); we publish the mean as the result and the minimum as the 1% low. Run-to-run variance is under 0.5%. Telemetry, power, temperature, utilisation, clocks, peak VRAM, is sampled at 1 Hz for the duration of every run. Where a model exceeds a card's VRAM we publish a hard won't-fit result rather than quietly dropping to a smaller quantisation. A card that can't run a model scores zero on it. Silently swapping precision to make a number appear would make every number on this site meaningless. All figures are single-GPU, single-stream, batch-size-1. That is the honest way to measure what one card does for one user, and it is deliberately not how a datacenter serves a model. Vendor and MLPerf figures use large batches across many GPUs and will be far higher. Neither is wrong, they answer different questions. Ours answers 'what will this card do for me'.