10 short social clips

A week of vertical video in one sitting, hooks written, frames generated, clips rendered.

The pipeline

  1. Write 10 hooks and captions, qwen3-32b
  2. Generate 10 opening frames, z-image-turbo
  3. Render 10 clips (~2s each), wan-2-2-5b

Fastest card where every stage is a real measurement: NVIDIA B300, 4.7 min for the whole job.

Every GPU, slowest job to fastest

GPUTotal timeComputeEnergyBasis
NVIDIA B3004.7 min3.1 min51.37 Whall 3 stages measured
NVIDIA B2005.6 min4.8 min69.91 Whall 3 stages measured
NVIDIA B1005.9 min5.9 minn/aanchored estimate (0/3 stages measured)
NVIDIA GH200 Grace Hopper6.3 min6.3 minn/aanchored estimate (0/3 stages measured)
NVIDIA H100 NVL6.7 min6.7 minn/aanchored estimate (0/3 stages measured)
NVIDIA H800 80GB6.7 min6.7 minn/aanchored estimate (0/3 stages measured)
NVIDIA H100 PCIe7.8 min7.8 minn/aanchored estimate (0/3 stages measured)
NVIDIA H100 80GB HBM38 min6.7 min73.89 Whall 3 stages measured
NVIDIA H20010 min6.3 min71.1 Whall 3 stages measured
NVIDIA RTX PRO 6000 Blackwell Server Edition10.8 min10.1 min88.63 Whall 3 stages measured
NVIDIA A800 80GB13.6 min13.6 minn/aanchored estimate (0/3 stages measured)
NVIDIA A100 40GB SXM413.7 min13.6 min0.74 Whanchored estimate (1/3 stages measured)
NVIDIA A100 40GB PCIe14.7 min14.7 minn/aanchored estimate (0/3 stages measured)
NVIDIA A100 80GB SXM415.1 min13.6 min87.69 Whall 3 stages measured
NVIDIA A100 80GB PCIe15.7 min14.7 min70.16 Whall 3 stages measured
NVIDIA RTX PRO 5000 Blackwell16.5 min15.8 min78.82 Whall 3 stages measured
NVIDIA L40S19.1 min18.1 min96.77 Whall 3 stages measured
NVIDIA GeForce RTX 409022.7 min20.5 min125.77 Whall 3 stages measured
NVIDIA RTX 5880 Ada Generation25.2 min25.2 minn/aanchored estimate (0/3 stages measured)
NVIDIA RTX 5000 Ada Generation27.9 min27.5 min97.41 Whall 3 stages measured
NVIDIA L4028.2 min24.1 min120.12 Whall 3 stages measured
NVIDIA RTX A550030.3 min30.3 minn/aanchored estimate (0/3 stages measured)
NVIDIA GeForce RTX 3090 Ti33.1 min32.7 min207.05 Whall 3 stages measured
NVIDIA A4033.4 min29.8 min145.56 Whall 3 stages measured
NVIDIA RTX PRO 4000 Blackwell34.2 min34 min78.69 Whall 3 stages measured
NVIDIA RTX A500037.2 min36.8 min129.95 Whall 3 stages measured
NVIDIA RTX 4500 Ada Generation38.9 min38.9 minn/aanchored estimate (0/3 stages measured)
NVIDIA GeForce RTX 309042.4 min40 min209.99 Whall 3 stages measured
NVIDIA A10G50.7 min48.6 min108.73 Whall 3 stages measured
AMD Radeon RX 7900 XTX58.7 min58.7 minn/aanchored estimate (0/3 stages measured)
NVIDIA L460.3 min59.2 min69.46 Whall 3 stages measured
AMD Radeon Pro W790062.5 min62.5 minn/aanchored estimate (0/3 stages measured)
AMD Radeon RX 7900 XT62.7 min62.7 minn/aanchored estimate (0/3 stages measured)
AMD Radeon Pro W780079.8 min79.8 minn/aanchored estimate (0/3 stages measured)
AMD Radeon Pro W680090 min90 minn/aanchored estimate (0/3 stages measured)

Cards that can't run this pipeline

Every one of these fails on the same kind of wall, a stage that will not fit in VRAM.

…and 18 more.

How these numbers are built

Each stage time is the quantity of work divided by that card's measured throughput for that model, from our own bench. The pipeline is assumed to run batched, every image, then every clip, so each model loads once. Model load time is added where we recorded it; our text-generation runs don't carry a load measurement yet, so pipelines with a language-model stage are slightly optimistic. Nothing here is a single timed run of the whole pipeline, and we don't present it as one.