40-product photo shoot

Generate the catalogue, then swap every background, the job that used to be a studio day.

The pipeline

  1. Generate 40 product shots, stable-diffusion-xl
  2. Swap the background on each, flux-1-kontext-dev

Fastest card where every stage is a real measurement: NVIDIA B300, 6.1 min for the whole job.

Every GPU, slowest job to fastest

GPUTotal timeComputeEnergyBasis
NVIDIA B3006.1 min5.2 min77.6 Whall 2 stages measured
NVIDIA B2008.4 min7.9 min113.44 Whall 2 stages measured
NVIDIA B1009.8 min9.8 minn/aanchored estimate (0/2 stages measured)
NVIDIA GH200 Grace Hopper10.2 min10.2 minn/aanchored estimate (0/2 stages measured)
NVIDIA H100 NVL10.5 min10.5 minn/aanchored estimate (0/2 stages measured)
NVIDIA H800 80GB10.5 min10.5 minn/aanchored estimate (0/2 stages measured)
NVIDIA H100 80GB HBM311.1 min10.5 min117.19 Whall 2 stages measured
NVIDIA H20011.2 min10.2 min115.44 Whall 2 stages measured
NVIDIA H100 PCIe12.2 min12.2 minn/aanchored estimate (0/2 stages measured)
NVIDIA RTX PRO 6000 Blackwell Server Edition15.2 min14.8 min133.59 Whall 2 stages measured
NVIDIA RTX PRO 6000 Blackwell Workstation Edition16.7 min14.4 min142.82 Whall 2 stages measured
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition18.5 min18.5 minn/aanchored estimate (0/2 stages measured)
NVIDIA A100 40GB SXM422.2 min22.2 min13.47 Whanchored estimate (1/2 stages measured)
NVIDIA A800 80GB22.5 min22.5 minn/aanchored estimate (0/2 stages measured)
NVIDIA A100 80GB SXM423.5 min22.5 min147.33 Whall 2 stages measured
NVIDIA A100 40GB PCIe24.1 min24.1 minn/aanchored estimate (0/2 stages measured)
NVIDIA A100 80GB PCIe24.8 min24.1 min118.29 Whall 2 stages measured
NVIDIA RTX PRO 5000 Blackwell25.1 min24.5 min122.39 Whall 2 stages measured
NVIDIA L40S26.7 min26 min146.58 Whall 2 stages measured
NVIDIA L4041.1 min39.5 min196.9 Whall 2 stages measured
NVIDIA RTX 6000 Ada Generation43 min42.6 min211.51 Whall 2 stages measured
NVIDIA RTX A600045.5 min42.1 min207.87 Whall 2 stages measured
NVIDIA RTX PRO 4500 Blackwell50.1 min48.6 min124.6 Whall 2 stages measured
NVIDIA RTX 5000 Ada Generation53.7 min53.6 min162.13 Whall 2 stages measured
NVIDIA RTX 5880 Ada Generation61.2 min61.2 minn/aanchored estimate (0/2 stages measured)
AMD Radeon Pro W790085.8 min85.8 minn/aanchored estimate (0/2 stages measured)
AMD Radeon Pro W78002 h 8 min2 h 8 minn/aanchored estimate (0/2 stages measured)
AMD Radeon Pro W68002 h 20 min2 h 20 minn/aanchored estimate (0/2 stages measured)

Cards that can't run this pipeline

Every one of these fails on the same kind of wall, a stage that will not fit in VRAM.

…and 28 more.

How these numbers are built

Each stage time is the quantity of work divided by that card's measured throughput for that model, from our own bench. The pipeline is assumed to run batched, every image, then every clip, so each model loads once. Model load time is added where we recorded it; our text-generation runs don't carry a load measurement yet, so pipelines with a language-model stage are slightly optimistic. Nothing here is a single timed run of the whole pipeline, and we don't present it as one.