AuraFlow v0.3 · 7 GPUs measured first-party · text-to-image · Updated October 2026
AuraFlow v0.3 on 7 GPUs, measured first-party: NVIDIA GeForce RTX 5090 leads at 4.35 images/min, L4 trails at 1.02 images/min, and it peaked at 20GB of VRAM.
Benchmarked weights: fal/AuraFlow-v0.3

4.35 images/min on AuraFlow v0.3, the ceiling. Measured on our bench. 32GB of VRAM, $1,999 at launch.

3.02 images/min on AuraFlow v0.3, fastest card you can buy at retail. Measured on our bench. 24GB of VRAM, $1,599 at launch.

1.5 images/min on AuraFlow v0.3, lowest launch price that still fits. Measured on our bench. 24GB of VRAM, $1,499 at launch.

1.67 images/min on AuraFlow v0.3, most speed per dollar. Measured on our bench. 24GB of VRAM, $1,999 at launch. That is 0.84 images/min per $1,000 of launch price.
What GPU Do You Need for AuraFlow v0.3?, images/min by GPU
Efficiency: images/min per 100W drawn
Power is the average pulled during the run, sampled at 1Hz. The fastest card is often not the one here, and for anything left running this is the number that shows up on the bill.
Value: images/min per $1,000 of MSRP
Launch price, not street price, so it ages. A speed leaderboard always crowns the most expensive card; this is the counterweight.
AuraFlow v0.3. Measured image generation speed by GPU
| GPU | Images/min | s per image | img/W·min | Avg power |
|---|---|---|---|---|
| NVIDIA GeForce RTX 5090 | 4.35 | 13.805 | 0.008 | 572.0 W |
| NVIDIA L40S | 3.58 | 16.74 | 0.01 | 347.5 W |
| NVIDIA GeForce RTX 4090 | 3.02 | 19.841 | 0.007 | 423.1 W |
| NVIDIA GeForce RTX 3090 Ti | 1.67 | 35.943 | 0.004 | 389.4 W |
| NVIDIA GeForce RTX 3090 | 1.5 | 39.997 | 0.005 | 328.8 W |
| NVIDIA A10G | 1.39 | 43.19 | 0.009 | 149.8 W |
| NVIDIA L4 | 1.02 | 58.96 | 0.014 | 72.0 W |
What the numbers show. Across 7 GPUs measured on our own bench, RTX 5090 is fastest at 4.35 images/min. The slowest, L4, manages 1.02, so the spread is 4.3x from top to bottom. L4 is the most efficient, 1.02 images/min at 72W.
How it compares. L40S: AuraFlow v0.3 3.58 images/min, FLUX.1 dev 3.84 (12B), ERNIE-Image Turbo 4.41 (8B), Stable Diffusion 3.5 Large 4.54 (8B), Qwen-Image 2.1 1.78 (7B). 3 of 4 beat AuraFlow v0.3 here.
Cost on a rented GPU. 1,000 images of AuraFlow v0.3: $1.36 on a RTX 3090 ($0.12/hr, 11.1 hours), $1.49 on a RTX 5090 ($0.39/hr, 3.8 hours, 1.1x the cost).
AuraFlow v0.3: cost per 1,000 images on rented GPUs
| GPU | Cheapest rate | Speed (images/min) | Cost per 1,000 images |
|---|---|---|---|
| NVIDIA GeForce RTX 3090 | $0.12/hr | 1.5 | $1.36 |
| NVIDIA GeForce RTX 5090 | $0.39/hr | 4.35 | $1.49 |
| NVIDIA GeForce RTX 4090 | $0.34/hr | 3.02 | $1.85 |
| NVIDIA GeForce RTX 3090 Ti | $0.27/hr | 1.67 | $2.69 |
| NVIDIA L40S | $0.79/hr | 3.58 | $3.68 |
| NVIDIA L4 | $0.44/hr | 1.02 | $7.19 |
Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.
Speed tiers for AuraFlow v0.3. under 6 images/min: 7 (RTX 5090, RTX 4090, RTX 3090 Ti). 30 images/min means two seconds or less per picture.
Time per image. AuraFlow v0.3 at 1024px and 28 steps: 13.8s per image on the RTX 5090, 58.8s on the L4. A batch of 100 takes 23 min on the fastest card and 98 min on the slowest.
VRAM for AuraFlow v0.3. Measured peak 18.7GB, so 24GB is the smallest common card size; smallest card it ran on: RTX 4090 (24GB).
Power on AuraFlow v0.3. Most efficient: L4, 72W, 1.18 kWh per 1,000 images. Hungriest: RTX 5090, 572W, 2.19 kWh. At $0.15/kWh: $0.18 per 1,000 images.
Fastest on AuraFlow v0.3: NVIDIA GeForce RTX 5090, 4.35 images/min. Cheapest consumer card that ran it: NVIDIA GeForce RTX 3090 ($1,499, 1.5 images/min). Cheapest to rent per job: NVIDIA GeForce RTX 3090, $1.36 per 1,000 images.
AuraFlow v0.3 at 1024px in diffusers, bf16, native precision with no offload, timed over three generations after a warmup, with power and VRAM sampled throughout. Image speed tracks tensor compute and architecture generation more than memory bandwidth, so the order here differs from our LLM boards.