Full codebase review

Feed a mid-size repository through a coding model and get notes back on every file.

The pipeline

  1. Read and annotate ~60k tokens of code, qwen2-5-coder-14b

Fastest card where every stage is a real measurement: NVIDIA B300, 6.3 min for the whole job.

Every GPU, slowest job to fastest

GPUTotal timeComputeEnergyBasis
NVIDIA H100 NVL5.9 min5.9 minn/aanchored estimate (0/1 stages measured)
NVIDIA B3006.3 min6.3 min38.61 Whall 1 stages measured
NVIDIA GH200 Grace Hopper6.6 min6.6 minn/aanchored estimate (0/1 stages measured)
NVIDIA B2006.6 min6.6 min44.06 Whall 1 stages measured
NVIDIA GeForce RTX 50906.7 min6.7 min13.03 Whall 1 stages measured
NVIDIA H2006.7 min6.7 min13.86 Whall 1 stages measured
NVIDIA RTX PRO 6000 Blackwell Workstation Edition6.8 min6.8 min27 Whall 1 stages measured
NVIDIA H100 80GB HBM36.9 min6.9 min26.8 Whall 1 stages measured
NVIDIA H800 80GB6.9 min6.9 minn/aanchored estimate (0/1 stages measured)
NVIDIA B1007 min7 minn/aanchored estimate (0/1 stages measured)
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition7.1 min7.1 minn/aanchored estimate (0/1 stages measured)
NVIDIA RTX PRO 6000 Blackwell Server Edition7.4 min7.4 min32.21 Whall 1 stages measured
NVIDIA RTX PRO 5000 Blackwell8.7 min8.7 min26.33 Whall 1 stages measured
NVIDIA GeForce RTX 409010.5 min10.5 min30.47 Whall 1 stages measured
NVIDIA A800 80GB11.2 min11.2 minn/aanchored estimate (0/1 stages measured)
NVIDIA A100 80GB SXM411.2 min11.2 min32.59 Whall 1 stages measured
NVIDIA A100 80GB PCIe11.3 min11.3 min33.67 Whall 1 stages measured
NVIDIA GeForce RTX 3090 Ti11.3 min11.3 min49.88 Whall 1 stages measured
NVIDIA RTX 6000 Ada Generation11.4 min11.4 min43.65 Whall 1 stages measured
NVIDIA H100 PCIe11.6 min11.6 minn/aanchored estimate (0/1 stages measured)
NVIDIA A100 40GB SXM411.7 min11.6 min35.27 Whall 1 stages measured
GeForce RTX 5070 Ti12 min12 min44.78 Whall 1 stages measured
GeForce RTX 508012.2 min12.2 min29.75 Whall 1 stages measured
NVIDIA RTX PRO 4500 Blackwell12.5 min12.5 min21.85 Whall 1 stages measured
NVIDIA GeForce RTX 309012.7 min12.7 min37.63 Whall 1 stages measured
NVIDIA GeForce RTX 3080 Ti12.8 min12.8 min64.96 Whall 1 stages measured
NVIDIA L40S13.5 min13.4 min49.6 Whall 1 stages measured
NVIDIA L4013.5 min13.5 min37.92 Whall 1 stages measured
NVIDIA A100 40GB PCIe14.1 min14.1 minn/aanchored estimate (0/1 stages measured)
GeForce RTX 4080 Super14.2 min14.2 min33.05 Whall 1 stages measured
NVIDIA GeForce RTX 408014.3 min14.3 min44.74 Whall 1 stages measured
NVIDIA RTX A550014.3 min14.3 minn/aanchored estimate (0/1 stages measured)
AMD Radeon RX 7900 XTX14.3 min14.3 minn/aanchored estimate (0/1 stages measured)
NVIDIA RTX A600014.6 min14.6 min41.58 Whall 1 stages measured
AMD Radeon Pro W790015 min15 minn/aanchored estimate (0/1 stages measured)
NVIDIA GeForce RTX 4070 Ti Super15.2 min15.2 min50.72 Whall 1 stages measured
NVIDIA RTX A500015.4 min15.4 min46.13 Whall 1 stages measured
NVIDIA RTX 5880 Ada Generation15.8 min15.8 minn/aanchored estimate (0/1 stages measured)
AMD Radeon RX 7900 XT16.1 min16.1 minn/aanchored estimate (0/1 stages measured)
NVIDIA RTX PRO 4000 Blackwell16.5 min16.5 min33.79 Whall 1 stages measured
NVIDIA Titan RTX16.7 min16.7 min42.48 Whall 1 stages measured
NVIDIA A4017.3 min17.3 min54.16 Whall 1 stages measured
NVIDIA RTX 5000 Ada Generation17.6 min17.6 min59.48 Whall 1 stages measured
NVIDIA TITAN V18.4 min18.4 minn/aanchored estimate (0/1 stages measured)
NVIDIA RTX A450018.4 min18.4 min42.18 Whall 1 stages measured
NVIDIA GeForce RTX 4070 Super19.5 min19.5 min46.03 Whall 1 stages measured
NVIDIA GeForce RTX 407019.7 min19.7 min43.13 Whall 1 stages measured
NVIDIA GeForce RTX 4070 Ti19.7 min19.7 minn/aanchored estimate (0/1 stages measured)
NVIDIA TITAN Xp19.9 min19.9 minn/aanchored estimate (0/1 stages measured)
AMD Radeon RX 9070 XT21.2 min21.2 minn/aanchored estimate (0/1 stages measured)
AMD Radeon RX 907021.2 min21.2 minn/aanchored estimate (0/1 stages measured)
NVIDIA A10G21.3 min21.2 min43.41 Whall 1 stages measured
NVIDIA Quadro RTX 800021.4 min21.4 min49.24 Whall 1 stages measured
NVIDIA Quadro RTX 6000 (Turing)21.5 min21.5 min62.2 Whall 1 stages measured
GeForce GTX 1080 Ti22.1 min22.1 minn/aanchored estimate (0/1 stages measured)
NVIDIA Quadro RTX 500022.2 min22.2 minn/aanchored estimate (0/1 stages measured)
GeForce RTX 5060 Ti22.2 min22.2 min45.98 Whall 1 stages measured
NVIDIA TITAN X (Pascal)22.3 min22.3 minn/aanchored estimate (0/1 stages measured)
AMD Radeon Pro W780022.7 min22.7 minn/aanchored estimate (0/1 stages measured)
AMD Radeon RX 6950 XT22.7 min22.7 minn/aanchored estimate (0/1 stages measured)
NVIDIA RTX 4500 Ada Generation22.9 min22.9 minn/aanchored estimate (0/1 stages measured)
AMD Radeon Pro W680022.9 min22.9 minn/aanchored estimate (0/1 stages measured)
AMD Radeon RX 6800 XT22.9 min22.9 minn/aanchored estimate (0/1 stages measured)
AMD Radeon RX 680022.9 min22.9 minn/aanchored estimate (0/1 stages measured)
AMD Radeon RX 6900 XT22.9 min22.9 minn/aanchored estimate (0/1 stages measured)
AMD Radeon RX 7800 XT22.9 min22.9 minn/aanchored estimate (0/1 stages measured)
NVIDIA RTX A400024.3 min24.3 min30.11 Whall 1 stages measured
AMD Radeon RX 7700 XT26.9 min26.9 minn/aanchored estimate (0/1 stages measured)
Intel Arc A770 Limited Edition27.3 min27.3 minn/aanchored estimate (0/1 stages measured)
NVIDIA RTX 4000 (Ada Generation)27.5 min27.5 min38.11 Whall 1 stages measured
NVIDIA GeForce RTX 306028.2 min28.2 min69.16 Whall 1 stages measured
NVIDIA GeForce RTX 4060 Ti 16GB32.1 min32.1 min67.32 Whall 1 stages measured
NVIDIA L436.4 min36.4 min37.51 Whall 1 stages measured
Intel Arc Pro A6036.8 min36.8 minn/aanchored estimate (0/1 stages measured)
AMD Radeon RX 670037.3 min37.3 minn/aanchored estimate (0/1 stages measured)
NVIDIA RTX 2000 Ada Generation42.8 min42.8 min32.76 Whall 1 stages measured
NVIDIA T450.7 min50.6 min51.37 Whall 1 stages measured

Cards that can't run this pipeline

Every one of these fails on the same kind of wall, a stage that will not fit in VRAM.

How these numbers are built

Each stage time is the quantity of work divided by that card's measured throughput for that model, from our own bench. The pipeline is assumed to run batched, every image, then every clip, so each model loads once. Model load time is added where we recorded it; our text-generation runs don't carry a load measurement yet, so pipelines with a language-model stage are slightly optimistic. Nothing here is a single timed run of the whole pipeline, and we don't present it as one.