FLUX.1 Schnell · 8 GPUs measured first-party · text-to-image · Updated July 2026

What GPU Do You Need for FLUX.1 Schnell?

FLUX.1 Schnell is the speed king of open image generation, a distilled 4-step version of Black Forest Labs' FLUX that produced the fastest text-to-image results in our database: 81 images per minute on a B200, 0.74 seconds per image. We measured it on 8 GPUs at bf16 with the full pipeline resident, which is where its ~37GB peak VRAM figure comes from.

Benchmarked weights: black-forest-labs/FLUX.1-schnell

81.41img/min
Fastest: NVIDIA B200
measured, 3-run llama-bench
~37GB
VRAM needed (measured peak)
GPU-independent, applies to every card
8
GPUs measured
same pinned harness
0.74s
Per image (fastest card)
1024px, native steps

What GPU Do You Need for FLUX.1 Schnell?, images/min, fastest 8

NVIDIA B200
81.41 Images/min
NVIDIA H200
60.51 Images/min
NVIDIA B300
59.54 Images/min
NVIDIA H100 80GB HBM3
58.36 Images/min
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
42.41 Images/min
NVIDIA A100 40GB SXM4
28.44 Images/min
NVIDIA A100 80GB SXM4
28.39 Images/min
NVIDIA L40S
26.24 Images/min

Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.

FLUX.1 Schnell. Measured image generation speed by GPU

NVIDIA B20081.41
NVIDIA H20060.51
NVIDIA B30059.54
NVIDIA H100 80GB HBM358.36
NVIDIA RTX PRO 6000 Blackwell Workstation Edition42.41
NVIDIA A100 40GB SXM428.44
NVIDIA A100 80GB SXM428.39
NVIDIA L40S26.24
GPUImages/mins per imageimg/W·minAvg power
NVIDIA B20081.410.740.106769.5 W
NVIDIA H20060.510.990.107566.8 W
NVIDIA B30059.541.010.096619.0 W
NVIDIA H100 80GB HBM358.361.030.095614.2 W
NVIDIA RTX PRO 6000 Blackwell Workstation Edition42.411.410.084502.7 W
NVIDIA A100 40GB SXM428.442.110.074383.5 W
NVIDIA A100 80GB SXM428.392.110.072393.8 W
NVIDIA L40S26.242.290.079330.8 W

Best-in-class speed, one known weakness. For raw image throughput, nothing open that we've measured touches Schnell, sub-second generation on Blackwell silicon changes what's practical: real-time preview loops, bulk asset generation, live creative tools. The caveat we'd flag from experience: text rendering inside images is its soft spot. The 4-step distillation that makes it fast costs it the legible signage, labels and typography that its bigger sibling FLUX.1 dev (and newer models like Z-Image Turbo) handle better. Composition prompt with no words in it? Schnell. Poster with a headline? Use dev.

About that VRAM number. Our ~37GB peak reflects the full bf16 pipeline held in memory, the way you'd serve it in production. That's datacenter or 48GB-workstation territory as tested. Consumer setups run Schnell every day via quantized weights and CPU offloading at lower speeds; we haven't measured those configurations yet, so this page won't guess at them. On rented silicon the math is sweet: 81 images/minute on a B200 means a thousand images costs about 12 minutes of GPU time.

Our verdict

FLUX.1 Schnell: 81 images/min, 0.74s per image, the fastest open text-to-image we've measured, with weak in-image text as the tradeoff of its 4-step distillation. As benchmarked (full bf16 pipeline, ~37GB) it's rented-GPU territory; its throughput-per-dollar there is unmatched.

FAQ

How fast is FLUX.1 Schnell really?
0.74 seconds per image on a B200, 81 images/minute sustained in our 10-batch protocol. The H200 and B300 land near 60/min. It's the fastest text-to-image result in our database by a wide margin.
Why does it need ~37GB of VRAM?
That's the full bf16 pipeline (transformer + text encoders + VAE) resident in memory, as we benchmarked it. Community setups fit it on consumer cards via quantization and offloading at reduced speed, configurations we haven't measured yet.
What's its weakness?
Text inside images. The 4-step speed distillation costs typography: signage, labels and headlines come out mangled far more often than with FLUX.1 dev. For text-heavy images, use dev or Z-Image Turbo and accept the slower generation.
Schnell or FLUX.1 dev?
Schnell for volume and iteration speed; dev for final quality and anything with words in it. A common pipeline: explore compositions on Schnell, re-render the keepers on dev.
Is renting cost-effective for it?
Extremely: at 81 images/minute, one B200 hour produces roughly 4,800 images. For bulk generation there is no cheaper open-model throughput in our data.