Z-Image · 7 GPUs measured first-party · text-to-image · Updated October 2026
Z-Image on 7 GPUs, measured first-party: NVIDIA H100 80GB HBM3 leads at 3.58 images/min, L4 trails at 0.37 images/min, and it peaked at 24GB of VRAM.
Benchmarked weights: Tongyi-MAI/Z-Image

3.58 images/min on Z-Image, the ceiling. Measured on our bench. 80GB of VRAM, $30,000 at launch.

1.72 images/min on Z-Image, fastest card you can buy at retail. Measured on our bench. 32GB of VRAM, $1,999 at launch.

0.63 images/min on Z-Image, lowest launch price that still fits. Measured on our bench. 24GB of VRAM, $1,499 at launch.

1.17 images/min on Z-Image, most speed per dollar. Measured on our bench. 24GB of VRAM, $1,599 at launch. That is 0.73 images/min per $1,000 of launch price.
What GPU Do You Need for Z-Image?, images/min by GPU
Efficiency: images/min per 100W drawn
Power is the average pulled during the run, sampled at 1Hz. The fastest card is often not the one here, and for anything left running this is the number that shows up on the bill.
Value: images/min per $1,000 of MSRP
Launch price, not street price, so it ages. A speed leaderboard always crowns the most expensive card; this is the counterweight.
Z-Image. Measured image generation speed by GPU
| GPU | Images/min | s per image | img/W·min | Avg power |
|---|---|---|---|---|
| NVIDIA H100 80GB HBM3 | 3.58 | 16.77 | 0.005 | 688.5 W |
| NVIDIA GeForce RTX 5090 | 1.72 | 34.837 | 0.003 | 524.9 W |
| NVIDIA L40S | 1.31 | 45.9 | 0.004 | 344.2 W |
| NVIDIA GeForce RTX 4090 | 1.17 | 51.291 | 0.003 | 398.3 W |
| NVIDIA GeForce RTX 3090 Ti | 0.7 | 85.197 | 0.002 | 394.2 W |
| NVIDIA GeForce RTX 3090 | 0.63 | 95.483 | 0.002 | 348.4 W |
| NVIDIA L4 | 0.37 | 160.59 | 0.005 | 73.4 W |
What the numbers show. Across 7 GPUs measured on our own bench, H100 80GB HBM3 is fastest at 3.58 images/min. The slowest, L4, manages 0.37, so the spread is 9.7x from top to bottom. Per dollar of launch price, RTX 5090 gives the most (0.9 images/min per $1,000).
About Z-Image. Z-Image: from Tongyi-MAI, 6.2B parameters, on Hugging Face since January 2026, Apache 2.0 licence. 193,076 downloads in the last 30 days and 3 community quantizations.
How it compares. L40S: Z-Image 1.31 images/min, Z-Image Turbo (1024px) 18.2 (6B), Qwen-Image 2.1 1.78 (7B), ERNIE-Image Turbo 4.41 (8B), Stable Diffusion 3.5 Large 4.54 (8B). All 4 beat Z-Image here.
Cost on a rented GPU. 1,000 images of Z-Image: $3.23 on a RTX 3090 ($0.12/hr, 26.5 hours), $9.94 on a H100 80GB HBM3 ($2.14/hr, 4.7 hours, 3.1x the cost).
Z-Image: cost per 1,000 images on rented GPUs
| GPU | Cheapest rate | Speed (images/min) | Cost per 1,000 images |
|---|---|---|---|
| NVIDIA GeForce RTX 3090 | $0.12/hr | 0.63 | $3.23 |
| NVIDIA GeForce RTX 5090 | $0.39/hr | 1.72 | $3.77 |
| NVIDIA GeForce RTX 4090 | $0.34/hr | 1.17 | $4.79 |
| NVIDIA GeForce RTX 3090 Ti | $0.27/hr | 0.7 | $6.43 |
| NVIDIA H100 80GB HBM3 | $2.14/hr | 3.58 | $9.94 |
| NVIDIA L40S | $0.79/hr | 1.31 | $10.05 |
| NVIDIA L4 | $0.44/hr | 0.37 | $19.82 |
Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.
Speed tiers for Z-Image. under 6 images/min: 7 (RTX 5090, RTX 4090, RTX 3090 Ti). 30 images/min means two seconds or less per picture.
Time per image. Z-Image at 1024px and 50 steps: 16.8s per image on the H100 80GB HBM3, 162.2s on the L4. A batch of 100 takes 28 min on the fastest card and 270 min on the slowest.
VRAM for Z-Image. Measured peak 22.3GB, so 24GB is the smallest common card size; smallest card it ran on: RTX 4090 (24GB).
Power on Z-Image. Most efficient: H100 80GB HBM3, 688W, 3.21 kWh per 1,000 images. At $0.15/kWh: $0.48 per 1,000 images.
Fastest on Z-Image: NVIDIA H100 80GB HBM3, 3.58 images/min. Best desktop card: NVIDIA GeForce RTX 5090, 1.72 images/min. Cheapest consumer card that ran it: NVIDIA GeForce RTX 3090 ($1,499, 0.63 images/min). Cheapest to rent per job: NVIDIA GeForce RTX 3090, $3.23 per 1,000 images.
Z-Image at 1024px in diffusers, bf16, native precision with no offload, timed over three generations after a warmup, with power and VRAM sampled throughout. Image speed tracks tensor compute and architecture generation more than memory bandwidth, so the order here differs from our LLM boards.