Script it, generate every keyframe, then render the clips. The full text-to-video pipeline end to end.
Fastest card where every stage is a real measurement: NVIDIA B200, 1.9 min for the whole job.
| GPU | Total time | Compute | Energy | Basis |
|---|---|---|---|---|
| NVIDIA B100 | 1.7 min | 1.7 min | n/a | anchored estimate (0/3 stages measured) |
| NVIDIA B200 | 1.9 min | 85 s | 16.03 Wh | all 3 stages measured |
| NVIDIA GH200 Grace Hopper | 2 min | 2 min | n/a | anchored estimate (0/3 stages measured) |
| NVIDIA H100 NVL | 2 min | 2 min | n/a | anchored estimate (0/3 stages measured) |
| NVIDIA H800 80GB | 2 min | 2 min | n/a | anchored estimate (0/3 stages measured) |
| NVIDIA B300 | 2.2 min | 87 s | 14.29 Wh | all 3 stages measured |
| NVIDIA H100 PCIe | 2.4 min | 2.4 min | n/a | anchored estimate (0/3 stages measured) |
| NVIDIA H100 80GB HBM3 | 2.5 min | 2 min | 17.46 Wh | all 3 stages measured |
| NVIDIA RTX PRO 6000 Blackwell Server Edition | 2.7 min | 2.4 min | 19.22 Wh | all 3 stages measured |
| NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition | 2.8 min | 2.8 min | n/a | anchored estimate (0/3 stages measured) |
| NVIDIA H200 | 3.9 min | 2 min | 18.03 Wh | all 3 stages measured |
| NVIDIA A800 80GB | 4 min | 4 min | n/a | anchored estimate (0/3 stages measured) |
| NVIDIA A100 40GB SXM4 | 4 min | 3.9 min | 6.39 Wh | anchored estimate (2/3 stages measured) |
| NVIDIA A100 40GB PCIe | 4.2 min | 4.2 min | n/a | anchored estimate (0/3 stages measured) |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 4.5 min | 2.2 min | 19.26 Wh | all 3 stages measured |
| NVIDIA A100 80GB PCIe | 4.6 min | 4.1 min | 18.6 Wh | all 3 stages measured |
| NVIDIA A100 80GB SXM4 | 4.7 min | 4 min | 23.61 Wh | all 3 stages measured |
| NVIDIA L40S | 4.9 min | 4.3 min | 22.52 Wh | all 3 stages measured |
| NVIDIA RTX PRO 5000 Blackwell | 6.7 min | 6.4 min | 20.33 Wh | all 3 stages measured |
| NVIDIA GeForce RTX 4090 | 7.9 min | 6.4 min | 30.54 Wh | all 3 stages measured |
| NVIDIA RTX 5880 Ada Generation | 8.6 min | 8.6 min | n/a | anchored estimate (0/3 stages measured) |
| NVIDIA RTX 5000 Ada Generation | 9.5 min | 9.3 min | 27.06 Wh | all 3 stages measured |
| NVIDIA GeForce RTX 3090 Ti | 9.9 min | 9.7 min | 50.99 Wh | all 3 stages measured |
| NVIDIA RTX PRO 4000 Blackwell | 10.2 min | 10.1 min | 20.63 Wh | all 3 stages measured |
| NVIDIA RTX PRO 4500 Blackwell | 10.6 min | 8.3 min | 18.99 Wh | all 3 stages measured |
| NVIDIA RTX A5000 | 11.8 min | 11.5 min | 36.17 Wh | all 3 stages measured |
| NVIDIA L40 | 13 min | 11.1 min | 36.37 Wh | all 3 stages measured |
| NVIDIA RTX A5500 | 13.1 min | 13.1 min | n/a | anchored estimate (0/3 stages measured) |
| NVIDIA RTX 4500 Ada Generation | 15.1 min | 15.1 min | n/a | anchored estimate (0/3 stages measured) |
| NVIDIA GeForce RTX 3090 | 15.2 min | 13.6 min | 57.1 Wh | all 3 stages measured |
| AMD Radeon RX 7900 XTX | 16.1 min | 16.1 min | n/a | anchored estimate (0/3 stages measured) |
| AMD Radeon Pro W7900 | 18.6 min | 18.6 min | n/a | anchored estimate (0/3 stages measured) |
| AMD Radeon RX 7900 XT | 20.7 min | 20.7 min | n/a | anchored estimate (0/3 stages measured) |
| AMD Radeon Pro W7800 | 24.3 min | 24.3 min | n/a | anchored estimate (0/3 stages measured) |
| AMD Radeon Pro W6800 | 28.8 min | 28.8 min | n/a | anchored estimate (0/3 stages measured) |
Every one of these fails on the same kind of wall, a stage that will not fit in VRAM.
…and 18 more.
Each stage time is the quantity of work divided by that card's measured throughput for that model, from our own bench. The pipeline is assumed to run batched, every image, then every clip, so each model loads once. Model load time is added where we recorded it; our text-generation runs don't carry a load measurement yet, so pipelines with a language-model stage are slightly optimistic. Nothing here is a single timed run of the whole pipeline, and we don't present it as one.