Cut every product off its background and upscale it for print. The job an online shop runs on every new stock delivery.
Fastest card where every stage is a real measurement: NVIDIA B300, 3.7 min for the whole job.
| GPU | Total time | Compute | Energy | Basis |
|---|---|---|---|---|
| NVIDIA B300 | 3.7 min | 3.6 min | 25.74 Wh | all 2 stages measured |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 4 min | 3.9 min | 23.56 Wh | all 2 stages measured |
| NVIDIA L40S | 6.2 min | 6.1 min | 24.48 Wh | all 2 stages measured |
| NVIDIA A100 80GB SXM4 | 7.3 min | 7.2 min | 21.66 Wh | all 2 stages measured |
| NVIDIA H200 | 8 min | 7.9 min | 27.94 Wh | all 2 stages measured |
| NVIDIA H100 80GB HBM3 | 8.1 min | 8 min | 26.87 Wh | all 2 stages measured |
| NVIDIA A10G | 9.6 min | 9.4 min | 22.58 Wh | all 2 stages measured |
| NVIDIA L4 | 12.1 min | 11.6 min | 13.61 Wh | all 2 stages measured |
| NVIDIA A100 40GB SXM4 | 12.2 min | 11.9 min | 22.9 Wh | all 2 stages measured |
| NVIDIA B200 | 17.2 min | 16.9 min | 79.19 Wh | all 2 stages measured |
| NVIDIA T4 | 20.3 min | 20.1 min | 21.65 Wh | all 2 stages measured |
Each stage time is the quantity of work divided by that card's measured throughput for that model, from our own bench. The pipeline is assumed to run batched, every image, then every clip, so each model loads once. Model load time is added where we recorded it; our text-generation runs don't carry a load measurement yet, so pipelines with a language-model stage are slightly optimistic. Nothing here is a single timed run of the whole pipeline, and we don't present it as one.