FLUX.1 Schnell · 8 GPUs measured first-party · text-to-image · Updated July 2026
FLUX.1 Schnell is the speed king of open image generation, a distilled 4-step version of Black Forest Labs' FLUX that produced the fastest text-to-image results in our database: 81 images per minute on a B200, 0.74 seconds per image. We measured it on 8 GPUs at bf16 with the full pipeline resident, which is where its ~37GB peak VRAM figure comes from.
Benchmarked weights: black-forest-labs/FLUX.1-schnell
What GPU Do You Need for FLUX.1 Schnell?, images/min, fastest 8
Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.
FLUX.1 Schnell. Measured image generation speed by GPU
| GPU | Images/min | s per image | img/W·min | Avg power |
|---|---|---|---|---|
| NVIDIA B200 | 81.41 | 0.74 | 0.106 | 769.5 W |
| NVIDIA H200 | 60.51 | 0.99 | 0.107 | 566.8 W |
| NVIDIA B300 | 59.54 | 1.01 | 0.096 | 619.0 W |
| NVIDIA H100 80GB HBM3 | 58.36 | 1.03 | 0.095 | 614.2 W |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 42.41 | 1.41 | 0.084 | 502.7 W |
| NVIDIA A100 40GB SXM4 | 28.44 | 2.11 | 0.074 | 383.5 W |
| NVIDIA A100 80GB SXM4 | 28.39 | 2.11 | 0.072 | 393.8 W |
| NVIDIA L40S | 26.24 | 2.29 | 0.079 | 330.8 W |
Best-in-class speed, one known weakness. For raw image throughput, nothing open that we've measured touches Schnell, sub-second generation on Blackwell silicon changes what's practical: real-time preview loops, bulk asset generation, live creative tools. The caveat we'd flag from experience: text rendering inside images is its soft spot. The 4-step distillation that makes it fast costs it the legible signage, labels and typography that its bigger sibling FLUX.1 dev (and newer models like Z-Image Turbo) handle better. Composition prompt with no words in it? Schnell. Poster with a headline? Use dev.
About that VRAM number. Our ~37GB peak reflects the full bf16 pipeline held in memory, the way you'd serve it in production. That's datacenter or 48GB-workstation territory as tested. Consumer setups run Schnell every day via quantized weights and CPU offloading at lower speeds; we haven't measured those configurations yet, so this page won't guess at them. On rented silicon the math is sweet: 81 images/minute on a B200 means a thousand images costs about 12 minutes of GPU time.
FLUX.1 Schnell: 81 images/min, 0.74s per image, the fastest open text-to-image we've measured, with weak in-image text as the tradeoff of its 4-step distillation. As benchmarked (full bf16 pipeline, ~37GB) it's rented-GPU territory; its throughput-per-dollar there is unmatched.