Z-Image Turbo · 34 cards measured first-party · Updated July 2026

How Fast Is Z-Image Turbo on Each GPU?

Z-Image Turbo is a modern distilled image model, 8 steps instead of 30, which sounds like it should be trivial. It isn't: it wants ~13GB minimum and ~23GB to run clean, so it gates hardware that handles SDXL without complaint.

Benchmarked weights: Tongyi-MAI/Z-Image-Turbo

Fastest we measured
NVIDIA B300

NVIDIA B300

5.29 it/s on Z-Image Turbo. Measured on our bench. 288GB of VRAM, 1400W board rating. AI Score 93.8/100 across our full 12-workload suite.

Pros
  • 5.29 it/s on Z-Image Turbo
  • 288GB, clears the Z-Image Turbo floor
  • Rentable by the hour rather than bought
Cons
  • 1400W board rating
  • Datacenter or workstation hardware, not a retail purchase

Best for: Z-Image Turbo work where you want the ceiling gone rather than the cheapest entry.

Runner-up
NVIDIA B200

NVIDIA B200

4.31 it/s on Z-Image Turbo. Measured on our bench. 192GB of VRAM, 1000W board rating. AI Score 78.0/100 across our full 12-workload suite.

Pros
  • 4.31 it/s on Z-Image Turbo
  • 192GB, clears the Z-Image Turbo floor
  • Rentable by the hour rather than bought
Cons
  • 1000W board rating
  • Datacenter or workstation hardware, not a retail purchase

Best for: Z-Image Turbo work where you want the ceiling gone rather than the cheapest entry.

Third
NVIDIA H200

NVIDIA H200

3.09 it/s on Z-Image Turbo. Measured on our bench. 141GB of VRAM, 700W board rating. AI Score 65.0/100 across our full 12-workload suite.

Pros
  • 3.09 it/s on Z-Image Turbo
  • 141GB, clears the Z-Image Turbo floor
  • Rentable by the hour rather than bought
Cons
  • 700W board rating
  • Datacenter or workstation hardware, not a retail purchase

Best for: Z-Image Turbo work where you want the ceiling gone rather than the cheapest entry.

5.29it/s
Fastest: NVIDIA B300
measured
43
Cards that run Z-Image Turbo
of 56 we have data for
13
Cards that can't run it at all
published as hard gates, not omissions
8817%
Fastest vs slowest that fits
5.29 vs 0.06 it/s

Compute-bound like all diffusion, but the step count changes the character of the benchmark. Fewer, heavier steps means less opportunity to hide latency, and the gap between architecture generations widens rather than narrows.

Z-Image Turbo, the 12 fastest cards we have data for

NVIDIA B300
5.29 it/s
NVIDIA B200
4.31 it/s
NVIDIA B100
3.45 it/s
NVIDIA GH200 Grace Hopper
3.09 it/s
NVIDIA H200
3.09 it/s
NVIDIA H100 NVL
2.95 it/s
NVIDIA H100 80GB HBM3
2.95 it/s
NVIDIA H800 80GB
2.95 it/s
NVIDIA H100 PCIe
2.55 it/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
2.08 it/s
NVIDIA RTX PRO 6000 Blackwell Server Edition
1.97 it/s
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
1.62 it/s

Single stream, batch size 1. 34 of the 56 cards on this page were measured first-party by us; the rest are anchored estimates against those measurements and are labelled in the table below.

Won't fit, Z-Image Turbo gates these cards outright

NVIDIA GeForce RTX 306012GB
NVIDIA GeForce RTX 3080 Ti12GB
NVIDIA GeForce RTX 407012GB
NVIDIA GeForce RTX 308010GB
NVIDIA GeForce RTX 2060 Super8GB
NVIDIA GeForce RTX 2070 SUPER8GB
NVIDIA GeForce RTX 20708GB
NVIDIA GeForce RTX 2080 Super8GB
NVIDIA GeForce RTX 2080 Founders Edition8GB
NVIDIA GeForce RTX 3060 Ti8GB
NVIDIA GeForce RTX 3070 Founders Edition8GB
GeForce RTX 40608GB
NVIDIA GeForce RTX 20606GB
GPUVRAMWhy it fails
NVIDIA GeForce RTX 306012GBrequires ~13GB VRAM
NVIDIA GeForce RTX 3080 Ti12GBrequires ~13GB VRAM
NVIDIA GeForce RTX 407012GBrequires ~13GB VRAM
NVIDIA GeForce RTX 308010GBrequires ~13GB VRAM
NVIDIA GeForce RTX 2060 Super8GBNeeds needs ~13GB VRAM
NVIDIA GeForce RTX 2070 SUPER8GBNeeds needs ~13GB VRAM
NVIDIA GeForce RTX 20708GBNeeds needs ~13GB VRAM
NVIDIA GeForce RTX 2080 Super8GBNeeds needs ~13GB VRAM
NVIDIA GeForce RTX 2080 Founders Edition8GBNeeds needs ~13GB VRAM
NVIDIA GeForce RTX 3060 Ti8GBrequires ~13GB VRAM
NVIDIA GeForce RTX 3070 Founders Edition8GBrequires ~13GB VRAM
GeForce RTX 40608GBrequires ~13GB VRAM
NVIDIA GeForce RTX 20606GBNeeds needs ~13GB VRAM

No driver update fixes a VRAM ceiling.

Full Z-Image Turbo leaderboard, every card that runs it

NVIDIA B3005.29 it/s
NVIDIA B2004.31 it/s
NVIDIA B1003.45 it/s
NVIDIA GH200 Grace Hopper3.09 it/s
NVIDIA H2003.09 it/s
NVIDIA H100 NVL2.95 it/s
NVIDIA H100 80GB HBM32.95 it/s
NVIDIA H800 80GB2.95 it/s
NVIDIA H100 PCIe2.55 it/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition2.08 it/s
NVIDIA RTX PRO 6000 Blackwell Server Edition1.97 it/s
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition1.62 it/s
NVIDIA GeForce RTX 50901.43 it/s
NVIDIA A100 40GB SXM41.33 it/s
NVIDIA A100 80GB SXM41.33 it/s
NVIDIA A800 80GB1.33 it/s
NVIDIA A100 40GB PCIe1.24 it/s
NVIDIA A100 80GB PCIe1.24 it/s
NVIDIA RTX PRO 5000 Blackwell1.21 it/s
NVIDIA L40S1.11 it/s
NVIDIA GeForce RTX 40900.99 it/s
NVIDIA RTX PRO 4500 Blackwell0.86 it/s
NVIDIA RTX 5000 Ada Generation0.79 it/s
NVIDIA L400.75 it/s
NVIDIA RTX A60000.72 it/s
NVIDIA RTX A55000.71 it/s
NVIDIA RTX 6000 Ada Generation0.68 it/s
NVIDIA RTX 5880 Ada Generation0.6 it/s
NVIDIA RTX PRO 4000 Blackwell0.59 it/s
NVIDIA RTX A50000.54 it/s
NVIDIA GeForce RTX 30900.51 it/s
NVIDIA A10G0.46 it/s
GeForce RTX 50800.36 it/s
AMD Radeon Pro W79000.34 it/s
NVIDIA GeForce RTX 40800.34 it/s
NVIDIA L40.33 it/s
GeForce RTX 5070 Ti0.24 it/s
NVIDIA RTX 4500 Ada Generation0.23 it/s
NVIDIA Quadro RTX 50000.2 it/s
NVIDIA GeForce RTX 4060 Ti0.18 it/s
AMD Radeon Pro W68000.17 it/s
AMD Radeon RX 6900 XT0.17 it/s
NVIDIA Quadro RTX 80000.06 it/s
GPUResultVRAMSource
NVIDIA B3005.29 it/s288GBMeasured
NVIDIA B2004.31 it/s192GBMeasured
NVIDIA B1003.45 it/s192GBEstimated
NVIDIA GH200 Grace Hopper3.09 it/s141GBEstimated
NVIDIA H2003.09 it/s141GBMeasured
NVIDIA H100 NVL2.95 it/s94GBEstimated
NVIDIA H100 80GB HBM32.95 it/s80GBMeasured
NVIDIA H800 80GB2.95 it/s80GBEstimated
NVIDIA H100 PCIe2.55 it/s80GBEstimated
NVIDIA RTX PRO 6000 Blackwell Workstation Edition2.08 it/s96GBMeasured
NVIDIA RTX PRO 6000 Blackwell Server Edition1.97 it/s96GBMeasured
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition1.62 it/s96GBEstimated
NVIDIA GeForce RTX 50901.43 it/s32GBMeasured
NVIDIA A100 40GB SXM41.33 it/s40GBEstimated
NVIDIA A100 80GB SXM41.33 it/s80GBMeasured
NVIDIA A800 80GB1.33 it/s80GBEstimated
NVIDIA A100 40GB PCIe1.24 it/s40GBEstimated
NVIDIA A100 80GB PCIe1.24 it/s80GBMeasured
NVIDIA RTX PRO 5000 Blackwell1.21 it/s48GBMeasured
NVIDIA L40S1.11 it/s48GBMeasured
NVIDIA GeForce RTX 40900.99 it/s24GBMeasured
NVIDIA RTX PRO 4500 Blackwell0.86 it/s32GBMeasured
NVIDIA RTX 5000 Ada Generation0.79 it/s32GBMeasured
NVIDIA L400.75 it/s48GBMeasured
NVIDIA RTX A60000.72 it/s48GBMeasured
NVIDIA RTX A55000.71 it/s24GBEstimated
NVIDIA RTX 6000 Ada Generation0.68 it/s48GBMeasured
NVIDIA RTX 5880 Ada Generation0.6 it/s48GBEstimated
NVIDIA RTX PRO 4000 Blackwell0.59 it/s24GBMeasured
NVIDIA RTX A50000.54 it/s24GBMeasured
NVIDIA GeForce RTX 30900.51 it/s24GBMeasured
NVIDIA A10G0.46 it/s24GBMeasured
GeForce RTX 50800.36 it/s16GBMeasured
AMD Radeon Pro W79000.34 it/s48GBEstimated
NVIDIA GeForce RTX 40800.34 it/s16GBMeasured
NVIDIA L40.33 it/s24GBMeasured
GeForce RTX 5070 Ti0.24 it/s16GBMeasured
NVIDIA RTX 4500 Ada Generation0.23 it/s24GBEstimated
NVIDIA Quadro RTX 50000.2 it/s16GBEstimated
NVIDIA GeForce RTX 4060 Ti0.18 it/s16GBMeasured
AMD Radeon Pro W68000.17 it/s32GBEstimated
AMD Radeon RX 6900 XT0.17 it/s16GBEstimated
NVIDIA Quadro RTX 80000.06 it/s48GBMeasured

Tap any column to sort. Measured = we rented and ran this card ourselves. Estimated = interpolated against our measured anchors, never blended silently.

Because this workload is tensor-compute bound, the ranking tracks architecture generation and tensor throughput rather than memory bandwidth, the reverse of our LLM charts. The same two cards can swap places entirely depending on which of these pages you're reading. That's the reason we run twelve workloads instead of publishing one score. A GPU isn't fast or slow. It's fast at some things and gated out of others, and which of those matters depends entirely on what you're actually going to run.

Our verdict

NVIDIA B300 tops our Z-Image Turbo leaderboard at 5.29 it/s (measured), 8817% of the way clear of the slowest card that still fits. But the number that decides most purchases isn't on the chart. It's the 13 cards that can't run Z-Image Turbo at all. This is a compute workload: buy architecture generation, not raw VRAM, as long as you clear the floor first.

FAQ

What is the fastest GPU for Z-Image Turbo?
NVIDIA B300, at 5.29 it/s on our bench, a first-party measurement. It carries 288GB of VRAM. Of the 56 cards we have Z-Image Turbo data for, 43 can run it at all.
How much VRAM do I need for Z-Image Turbo?
~13GB minimum, ~23GB for the full path. Cards that offload here take a brutal penalty rather than a graceful one.
Why does the Z-Image Turbo ranking look different from your other benchmarks?
Because this workload is tensor-compute bound, the ranking tracks architecture generation and tensor throughput rather than memory bandwidth, the reverse of our LLM charts. The same two cards can swap places entirely depending on which of these pages you're reading. That's why we publish twelve separate workloads rather than one blended score, the ordering genuinely changes depending on the job.
Are these Z-Image Turbo numbers measured or estimated?
Both, and every row says which. 34 of the 56 cards here were rented and run by us on the same harness. The remainder are anchored estimates interpolated per workload against those measurements. We never blend the two silently, if a row says Estimated, we have not run that card.
Can I rent a GPU to run Z-Image Turbo instead of buying one?
Yes, and for the cards at the top of this leaderboard it's the only realistic option, most of them have no retail channel at all. It's also how we got these numbers: we rented the hardware by the hour rather than buying it. That's worth considering before you spend on a card to find out whether it's fast enough.
Why publish cards that can't run Z-Image Turbo?
Because it's the most useful thing we know. A card that can't load a model doesn't run it slowly, it doesn't run it. Most benchmark sites leave that as a blank cell or quietly drop to a smaller quantisation to produce a number. We publish it as a hard gate and score it zero, because 'this card cannot do the thing you want' is the answer to the question you were actually asking.

How we test

Every ranking on this page comes from our own benchmark runs, not vendor claims. Cards marked Measured were rented and run by us; cards marked Estimated are interpolated per workload against those measured anchors and are labelled on every row, we never blend the two silently. LLMs run on llama.cpp (llama-bench) at Q4_K_M with -p 512 -n 128. Diffusion and video run on diffusers/ComfyUI at BF16, with SDXL at FP16. Each workload gets a warmup pass plus multiple timed runs (5 for small LLMs, 3 for large models and images, 2 for video); we publish the mean as the result and the minimum as the 1% low. Run-to-run variance is under 0.5%. Telemetry, power, temperature, utilisation, clocks, peak VRAM, is sampled at 1 Hz for the duration of every run. Where a model exceeds a card's VRAM we publish a hard won't-fit result rather than quietly dropping to a smaller quantisation. A card that can't run a model scores zero on it. Silently swapping precision to make a number appear would make every number on this site meaningless. All figures are single-GPU, single-stream, batch-size-1. That is the honest way to measure what one card does for one user, and it is deliberately not how a datacenter serves a model. Vendor and MLPerf figures use large batches across many GPUs and will be far higher. Neither is wrong, they answer different questions. Ours answers 'what will this card do for me'.