Wan 2.2 5B · 32 cards measured first-party · Updated July 2026
Wan 2.2 TI2V-5B generates a 49-frame clip at 1280×704, real 720p AI video. It needs ~18GB of VRAM minimum and ~38GB of system RAM, and it is the most honest test in our suite of whether a machine can do modern video generation or just talk about it.
Benchmarked weights: Wan-AI/Wan2.2-TI2V-5B-Diffusers

2.94 frames/s on Wan 2.2 5B. Measured on our bench. 288GB of VRAM, 1400W board rating. AI Score 93.8/100 across our full 12-workload suite.
Best for: Wan 2.2 5B work where you want the ceiling gone rather than the cheapest entry.

1.88 frames/s on Wan 2.2 5B. Measured on our bench. 192GB of VRAM, 1000W board rating. AI Score 78.0/100 across our full 12-workload suite.
Best for: Wan 2.2 5B work where you want the ceiling gone rather than the cheapest entry.

1.41 frames/s on Wan 2.2 5B. Measured on our bench. 141GB of VRAM, 700W board rating. AI Score 65.0/100 across our full 12-workload suite.
Best for: Wan 2.2 5B work where you want the ceiling gone rather than the cheapest entry.
Compute-bound and heavy. On our measured B300 this workload ran at 95.9% utilisation, drew over a kilowatt, and produced the hottest temperature we logged across all 21 workloads. Video is the one thing that reliably makes a datacenter GPU behave like a datacenter GPU.
Wan 2.2 5B, the 12 fastest cards we have data for
Single stream, batch size 1. 32 of the 54 cards on this page were measured first-party by us; the rest are anchored estimates against those measurements and are labelled in the table below.
Won't fit, Wan 2.2 5B gates these cards outright
| GPU | VRAM | Why it fails |
|---|---|---|
| NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition | 96GB | Estimated: won't fit or no sibling datum |
| AMD Radeon RX 6900 XT | 16GB | Needs needs ~18GB VRAM |
| NVIDIA GeForce RTX 4060 Ti | 16GB | requires ~18GB VRAM |
| NVIDIA GeForce RTX 4080 | 16GB | requires ~18GB VRAM |
| GeForce RTX 5070 Ti | 16GB | requires ~18GB VRAM |
| GeForce RTX 5080 | 16GB | requires ~18GB VRAM |
| NVIDIA Quadro RTX 5000 | 16GB | Needs needs ~18GB VRAM |
| NVIDIA RTX 2000 Ada Generation | 16GB | requires ~18GB VRAM |
| NVIDIA RTX A4000 | 16GB | requires ~18GB VRAM |
| NVIDIA GeForce RTX 3060 | 12GB | requires ~18GB VRAM |
| NVIDIA GeForce RTX 3080 Ti | 12GB | requires ~18GB VRAM |
| NVIDIA GeForce RTX 4070 | 12GB | requires ~18GB VRAM |
| NVIDIA GeForce RTX 3080 | 10GB | requires ~18GB VRAM |
| NVIDIA GeForce RTX 2060 Super | 8GB | Needs needs ~18GB VRAM |
Showing 14 of 22. No driver update fixes a VRAM ceiling.
Full Wan 2.2 5B leaderboard, every card that runs it
| GPU | Result | VRAM | Source |
|---|---|---|---|
| NVIDIA B300 | 2.94 frames/s | 288GB | Measured |
| NVIDIA B200 | 1.88 frames/s | 192GB | Measured |
| NVIDIA B100 | 1.5 frames/s | 192GB | Estimated |
| NVIDIA GH200 Grace Hopper | 1.41 frames/s | 141GB | Estimated |
| NVIDIA H200 | 1.41 frames/s | 141GB | Measured |
| NVIDIA H100 NVL | 1.33 frames/s | 94GB | Estimated |
| NVIDIA H100 80GB HBM3 | 1.33 frames/s | 80GB | Measured |
| NVIDIA H800 80GB | 1.33 frames/s | 80GB | Estimated |
| NVIDIA H100 PCIe | 1.15 frames/s | 80GB | Estimated |
| NVIDIA RTX PRO 6000 Blackwell Server Edition | 0.88 frames/s | 96GB | Measured |
| NVIDIA A100 40GB SXM4 | 0.66 frames/s | 40GB | Estimated |
| NVIDIA A100 80GB SXM4 | 0.66 frames/s | 80GB | Measured |
| NVIDIA A800 80GB | 0.66 frames/s | 80GB | Estimated |
| NVIDIA A100 40GB PCIe | 0.61 frames/s | 40GB | Estimated |
| NVIDIA A100 80GB PCIe | 0.61 frames/s | 80GB | Measured |
| NVIDIA RTX PRO 5000 Blackwell | 0.56 frames/s | 48GB | Measured |
| NVIDIA L40S | 0.49 frames/s | 48GB | Measured |
| NVIDIA GeForce RTX 4090 | 0.43 frames/s | 24GB | Measured |
| NVIDIA L40 | 0.37 frames/s | 48GB | Measured |
| NVIDIA RTX 5880 Ada Generation | 0.36 frames/s | 48GB | Estimated |
| NVIDIA RTX 5000 Ada Generation | 0.32 frames/s | 32GB | Measured |
| NVIDIA RTX A5500 | 0.29 frames/s | 24GB | Estimated |
| NVIDIA RTX PRO 4000 Blackwell | 0.26 frames/s | 24GB | Measured |
| NVIDIA RTX 4500 Ada Generation | 0.25 frames/s | 24GB | Estimated |
| NVIDIA RTX A5000 | 0.24 frames/s | 24GB | Measured |
| NVIDIA GeForce RTX 3090 | 0.22 frames/s | 24GB | Measured |
| NVIDIA RTX A4500 | 0.21 frames/s | 20GB | Measured |
| NVIDIA A10G | 0.18 frames/s | 24GB | Measured |
| NVIDIA RTX 4000 (Ada Generation) | 0.18 frames/s | 20GB | Measured |
| NVIDIA L4 | 0.15 frames/s | 24GB | Measured |
| AMD Radeon Pro W7900 | 0.14 frames/s | 48GB | Estimated |
| AMD Radeon Pro W6800 | 0.1 frames/s | 32GB | Estimated |
Tap any column to sort. Measured = we rented and ran this card ourselves. Estimated = interpolated against our measured anchors, never blended silently.
Because this workload is tensor-compute bound, the ranking tracks architecture generation and tensor throughput rather than memory bandwidth, the reverse of our LLM charts. The same two cards can swap places entirely depending on which of these pages you're reading. That's the reason we run twelve workloads instead of publishing one score. A GPU isn't fast or slow. It's fast at some things and gated out of others, and which of those matters depends entirely on what you're actually going to run.
NVIDIA B300 tops our Wan 2.2 5B leaderboard at 2.94 frames/s (measured), 2940% of the way clear of the slowest card that still fits. But the number that decides most purchases isn't on the chart. It's the 22 cards that can't run Wan 2.2 5B at all. This is a compute workload: buy architecture generation, not raw VRAM, as long as you clear the floor first.
Every ranking on this page comes from our own benchmark runs, not vendor claims. Cards marked Measured were rented and run by us; cards marked Estimated are interpolated per workload against those measured anchors and are labelled on every row, we never blend the two silently. LLMs run on llama.cpp (llama-bench) at Q4_K_M with -p 512 -n 128. Diffusion and video run on diffusers/ComfyUI at BF16, with SDXL at FP16. Each workload gets a warmup pass plus multiple timed runs (5 for small LLMs, 3 for large models and images, 2 for video); we publish the mean as the result and the minimum as the 1% low. Run-to-run variance is under 0.5%. Telemetry, power, temperature, utilisation, clocks, peak VRAM, is sampled at 1 Hz for the duration of every run. Where a model exceeds a card's VRAM we publish a hard won't-fit result rather than quietly dropping to a smaller quantisation. A card that can't run a model scores zero on it. Silently swapping precision to make a number appear would make every number on this site meaningless. All figures are single-GPU, single-stream, batch-size-1. That is the honest way to measure what one card does for one user, and it is deliberately not how a datacenter serves a model. Vendor and MLPerf figures use large batches across many GPUs and will be far higher. Neither is wrong, they answer different questions. Ours answers 'what will this card do for me'.