A month of written content in one run, on the biggest open model that fits your card.
Fastest card where every stage is a real measurement: NVIDIA B300, 9.7 min for the whole job.
| GPU | Total time | Compute | Energy | Basis |
|---|---|---|---|---|
| NVIDIA H100 NVL | 9.7 min | 9.7 min | n/a | anchored estimate (0/1 stages measured) |
| NVIDIA B300 | 9.7 min | 9.7 min | 61.26 Wh | all 1 stages measured |
| NVIDIA B200 | 10.5 min | 10.5 min | 67.6 Wh | all 1 stages measured |
| NVIDIA GH200 Grace Hopper | 10.7 min | 10.7 min | n/a | anchored estimate (0/1 stages measured) |
| NVIDIA H200 | 10.9 min | 10.9 min | 39.82 Wh | all 1 stages measured |
| NVIDIA B100 | 11 min | 11 min | n/a | anchored estimate (0/1 stages measured) |
| NVIDIA H100 80GB HBM3 | 11.4 min | 11.4 min | 55.51 Wh | all 1 stages measured |
| NVIDIA H800 80GB | 11.4 min | 11.4 min | n/a | anchored estimate (0/1 stages measured) |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 13.4 min | 13.4 min | 42.96 Wh | all 1 stages measured |
| NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition | 14.1 min | 14.1 min | n/a | anchored estimate (0/1 stages measured) |
| NVIDIA RTX PRO 6000 Blackwell Server Edition | 14.5 min | 14.5 min | 60.83 Wh | all 1 stages measured |
| NVIDIA RTX PRO 5000 Blackwell | 17.8 min | 17.8 min | 46.68 Wh | all 1 stages measured |
| NVIDIA H100 PCIe | 19 min | 19 min | n/a | anchored estimate (0/1 stages measured) |
| NVIDIA A100 80GB SXM4 | 19.1 min | 19.1 min | 55.72 Wh | all 1 stages measured |
| NVIDIA A800 80GB | 19.1 min | 19.1 min | n/a | anchored estimate (0/1 stages measured) |
| NVIDIA A100 80GB PCIe | 20.4 min | 20.4 min | 55.11 Wh | all 1 stages measured |
| NVIDIA RTX 6000 Ada Generation | 25.4 min | 25.4 min | 89.53 Wh | all 1 stages measured |
| NVIDIA L40S | 28.5 min | 28.3 min | 122.26 Wh | all 1 stages measured |
| NVIDIA RTX A6000 | 29.7 min | 29.7 min | 56.52 Wh | all 1 stages measured |
| AMD Radeon Pro W7900 | 33.1 min | 33.1 min | n/a | anchored estimate (0/1 stages measured) |
| NVIDIA RTX 5880 Ada Generation | 34.6 min | 34.6 min | n/a | anchored estimate (0/1 stages measured) |
Every one of these fails on the same kind of wall, a stage that will not fit in VRAM.
…and 38 more.
Each stage time is the quantity of work divided by that card's measured throughput for that model, from our own bench. The pipeline is assumed to run batched, every image, then every clip, so each model loads once. Model load time is added where we recorded it; our text-generation runs don't carry a load measurement yet, so pipelines with a language-model stage are slightly optimistic. Nothing here is a single timed run of the whole pipeline, and we don't present it as one.