20 long-form articles

A month of written content in one run, on the biggest open model that fits your card.

The pipeline

  1. Draft 20 articles (~1,400 tokens each), llama-3-3-70b

Fastest card where every stage is a real measurement: NVIDIA B300, 9.7 min for the whole job.

Every GPU, slowest job to fastest

GPUTotal timeComputeEnergyBasis
NVIDIA H100 NVL9.7 min9.7 minn/aanchored estimate (0/1 stages measured)
NVIDIA B3009.7 min9.7 min61.26 Whall 1 stages measured
NVIDIA B20010.5 min10.5 min67.6 Whall 1 stages measured
NVIDIA GH200 Grace Hopper10.7 min10.7 minn/aanchored estimate (0/1 stages measured)
NVIDIA H20010.9 min10.9 min39.82 Whall 1 stages measured
NVIDIA B10011 min11 minn/aanchored estimate (0/1 stages measured)
NVIDIA H100 80GB HBM311.4 min11.4 min55.51 Whall 1 stages measured
NVIDIA H800 80GB11.4 min11.4 minn/aanchored estimate (0/1 stages measured)
NVIDIA RTX PRO 6000 Blackwell Workstation Edition13.4 min13.4 min42.96 Whall 1 stages measured
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition14.1 min14.1 minn/aanchored estimate (0/1 stages measured)
NVIDIA RTX PRO 6000 Blackwell Server Edition14.5 min14.5 min60.83 Whall 1 stages measured
NVIDIA RTX PRO 5000 Blackwell17.8 min17.8 min46.68 Whall 1 stages measured
NVIDIA H100 PCIe19 min19 minn/aanchored estimate (0/1 stages measured)
NVIDIA A100 80GB SXM419.1 min19.1 min55.72 Whall 1 stages measured
NVIDIA A800 80GB19.1 min19.1 minn/aanchored estimate (0/1 stages measured)
NVIDIA A100 80GB PCIe20.4 min20.4 min55.11 Whall 1 stages measured
NVIDIA RTX 6000 Ada Generation25.4 min25.4 min89.53 Whall 1 stages measured
NVIDIA L40S28.5 min28.3 min122.26 Whall 1 stages measured
NVIDIA RTX A600029.7 min29.7 min56.52 Whall 1 stages measured
AMD Radeon Pro W790033.1 min33.1 minn/aanchored estimate (0/1 stages measured)
NVIDIA RTX 5880 Ada Generation34.6 min34.6 minn/aanchored estimate (0/1 stages measured)

Cards that can't run this pipeline

Every one of these fails on the same kind of wall, a stage that will not fit in VRAM.

…and 38 more.

How these numbers are built

Each stage time is the quantity of work divided by that card's measured throughput for that model, from our own bench. The pipeline is assumed to run batched, every image, then every clip, so each model loads once. Model load time is added where we recorded it; our text-generation runs don't carry a load measurement yet, so pipelines with a language-model stage are slightly optimistic. Nothing here is a single timed run of the whole pipeline, and we don't present it as one.