100-photo restore and enlarge

Repair a shoebox of old photos, then enlarge them enough to print.

The pipeline

  1. Restore and retouch 100 photos, flux-1-kontext-dev
  2. Upscale each 4x for print, swin2sr-4x

Fastest card where every stage is a real measurement: NVIDIA B300, 11.9 min for the whole job.

Every GPU, slowest job to fastest

GPUTotal timeComputeEnergyBasis
NVIDIA B30011.9 min11.4 min178.61 Whall 2 stages measured
NVIDIA B20026.4 min25.8 min296.86 Whall 2 stages measured
NVIDIA H20027.2 min26.6 min275.33 Whall 2 stages measured
NVIDIA H100 80GB HBM327.8 min27.4 min280.28 Whall 2 stages measured
NVIDIA RTX PRO 6000 Blackwell Workstation Edition35.3 min34.3 min334.41 Whall 2 stages measured
NVIDIA A100 80GB SXM454.3 min53.7 min341.34 Whall 2 stages measured
NVIDIA A100 40GB SXM456 min56 min11.25 Whanchored estimate (1/2 stages measured)
NVIDIA L40S62.5 min62 min346.58 Whall 2 stages measured

Cards that can't run this pipeline

Every one of these fails on the same kind of wall, a stage that will not fit in VRAM.

…and 31 more.

How these numbers are built

Each stage time is the quantity of work divided by that card's measured throughput for that model, from our own bench. The pipeline is assumed to run batched, every image, then every clip, so each model loads once. Model load time is added where we recorded it; our text-generation runs don't carry a load measurement yet, so pipelines with a language-model stage are slightly optimistic. Nothing here is a single timed run of the whole pipeline, and we don't present it as one.