200-product catalogue cutout

Cut every product off its background and upscale it for print. The job an online shop runs on every new stock delivery.

The pipeline

  1. Remove the background on 200 shots, birefnet
  2. Upscale each 4x, swin2sr-4x

Fastest card where every stage is a real measurement: NVIDIA B300, 3.7 min for the whole job.

Every GPU, slowest job to fastest

GPUTotal timeComputeEnergyBasis
NVIDIA B3003.7 min3.6 min25.74 Whall 2 stages measured
NVIDIA RTX PRO 6000 Blackwell Workstation Edition4 min3.9 min23.56 Whall 2 stages measured
NVIDIA L40S6.2 min6.1 min24.48 Whall 2 stages measured
NVIDIA A100 80GB SXM47.3 min7.2 min21.66 Whall 2 stages measured
NVIDIA H2008 min7.9 min27.94 Whall 2 stages measured
NVIDIA H100 80GB HBM38.1 min8 min26.87 Whall 2 stages measured
NVIDIA A10G9.6 min9.4 min22.58 Whall 2 stages measured
NVIDIA L412.1 min11.6 min13.61 Whall 2 stages measured
NVIDIA A100 40GB SXM412.2 min11.9 min22.9 Whall 2 stages measured
NVIDIA B20017.2 min16.9 min79.19 Whall 2 stages measured
NVIDIA T420.3 min20.1 min21.65 Whall 2 stages measured

How these numbers are built

Each stage time is the quantity of work divided by that card's measured throughput for that model, from our own bench. The pipeline is assumed to run batched, every image, then every clip, so each model loads once. Model load time is added where we recorded it; our text-generation runs don't carry a load measurement yet, so pipelines with a language-model stage are slightly optimistic. Nothing here is a single timed run of the whole pipeline, and we don't present it as one.