Image-to-video · 8 models · 15 GPUs measured first-party · Updated October 2026
Which graphics card to use for image-to-video, from first-party measurements of Stable Video Diffusion, CogVideoX-5B I2V, LTX-Video (image to video) and more on 15 GPUs.

4.88 clips/min on Stable Video Diffusion, the ceiling. Measured on our bench. 80GB of VRAM, $30,000 at launch.

2.84 clips/min on Stable Video Diffusion, fastest card you can buy at retail. Measured on our bench. 32GB of VRAM, $1,999 at launch.

0.43 clips/min on Stable Video Diffusion, lowest launch price that still fits. Measured on our bench. 12GB of VRAM, $329 at launch.

1.4 clips/min on Stable Video Diffusion, most speed per dollar. Measured on our bench. 16GB of VRAM, $749 at launch. That is 1.87 clips/min per $1,000 of launch price.
Image-to-video models take one still and animate it into a short clip: a few seconds of motion from a photo, a product shot or a generated frame. They are the heaviest consumer AI job we measure: a single clip is billions of operations per frame, repeated across dozens of denoising steps, and the bigger models need more VRAM than any consumer card has.
We measured 8 models for image-to-video on 15 GPUs. Speed is clips per minute, each clip being the model's standard length at the stated resolution. Every number below is a first-party run on our own harness; cards absent from a model's chart have not been run on it yet.
Stable Video Diffusion: clips/min by GPU
CogVideoX-5B I2V: clips/min by GPU
LTX-Video (image to video): clips/min by GPU
Stable Video Diffusion XT: clips/min by GPU
Wan 2.2 TI2V-5B (image to video): clips/min by GPU
Wan 2.1 I2V 14B (480p): clips/min by GPU
Cosmos-Predict2 2B Video2World: clips/min by GPU
Wan 2.2 I2V A14B: clips/min by GPU
Which models fit which card, for image-to-video
| Model | VRAM used | 8GB card | 12GB card | 16GB card | 24GB card | 32GB card | Licence |
|---|---|---|---|---|---|---|---|
| Stable Video Diffusion | 11.1GB | No | Yes | Yes | Yes | Yes | Stability Community (revenue cap) |
| CogVideoX-5B I2V | 15.9GB | No | No | No | Yes | Yes | CogVideoX licence (registration) |
| LTX-Video (image to video) | 12.4GB | No | No | Yes | Yes | Yes | LTX Community (revenue cap) |
| Stable Video Diffusion XT | 13.3GB | No | No | Yes | Yes | Yes | Stability Community (revenue cap) |
| Wan 2.2 TI2V-5B (image to video) | 15.3GB | No | No | Yes | Yes | Yes | Apache-2.0 |
| Wan 2.1 I2V 14B (480p) | 51.9GB | No | No | No | No | No | Apache-2.0 |
| Cosmos-Predict2 2B Video2World | 20.8GB | No | No | No | Yes | Yes | NVIDIA Open Model |
| Wan 2.2 I2V A14B | 74.1GB | No | No | No | No | No | Apache-2.0 |
From the lowest VRAM peak we measured for each model, plus 5% headroom. 'No' means it did not fit in that much memory at our settings, not that no setting ever could.
Every GPU x every image-to-video model (clips/min)
| GPU | SVD | CogVideoX-5B | LTX-Video | SVD-XT | Wan 2.2 5B | Wan 2.1 14B | Cosmos-Predict2 2B | Wan 2.2 A14B |
|---|---|---|---|---|---|---|---|---|
| NVIDIA H100 80GB HBM3 | 4.88 | 0.88 | 10.62 | 2.88 | 4.13 | 0.4 | — | 0.4 |
| NVIDIA GeForce RTX 5090 | 2.84 | 0.45 | 6.15 | 1.34 | 2.28 | — | — | — |
| NVIDIA GeForce RTX 4090 | 2.23 | 0.33 | 4.19 | 1.21 | 1.12 | — | 0.06 | — |
| NVIDIA L40S | 2.19 | 0.38 | 5.09 | 1.17 | 1.61 | — | — | — |
| GeForce RTX 5080 | 1.67 | 0.26 | 2.39 | 0.94 | 0.96 | — | — | — |
| NVIDIA GeForce RTX 4080 | 1.44 | 0.22 | 1.87 | 0.81 | 0.8 | — | — | — |
| GeForce RTX 5070 Ti | 1.4 | 0.21 | 1.82 | 0.77 | 0.77 | — | — | — |
| NVIDIA GeForce RTX 3090 | 1.06 | 0.15 | 2.17 | 0.49 | 0.6 | — | — | — |
| NVIDIA GeForce RTX 4070 Super | 1.03 | — | — | — | — | — | — | — |
| GeForce RTX 5060 Ti | 0.67 | 0.11 | 1.07 | 0.41 | 0.42 | — | — | — |
| NVIDIA GeForce RTX 4060 Ti 16GB | 0.64 | 0.1 | 0.96 | 0.36 | 0.37 | — | — | — |
| NVIDIA L4 | 0.63 | — | 1.41 | — | — | — | — | — |
| NVIDIA GeForce RTX 3060 | 0.43 | — | — | — | — | — | — | — |
| NVIDIA A100 80GB SXM4 | — | 0.47 | — | 1.54 | 2.16 | 0.2 | — | 0.21 |
| NVIDIA H200 | — | — | — | — | — | 0.42 | 0.14 | — |
— = not measured on that card yet.
What the numbers show.
Stable Video Diffusion: fastest on the NVIDIA H100 80GB HBM3 at 4.88 clips/min, 11.42x the slowest card we measured (NVIDIA GeForce RTX 3060); best consumer result NVIDIA GeForce RTX 5090 at 2.84; it used about 11.1GB of VRAM.
CogVideoX-5B I2V: fastest on the NVIDIA H100 80GB HBM3 at 0.88 clips/min, 8.88x the slowest card we measured (NVIDIA GeForce RTX 4060 Ti 16GB); best consumer result NVIDIA GeForce RTX 5090 at 0.45; it used about 15.9GB of VRAM.
LTX-Video (image to video): fastest on the NVIDIA H100 80GB HBM3 at 10.62 clips/min, 11.04x the slowest card we measured (NVIDIA GeForce RTX 4060 Ti 16GB); best consumer result NVIDIA GeForce RTX 5090 at 6.15; it used about 12.4GB of VRAM.
Stable Video Diffusion XT: fastest on the NVIDIA H100 80GB HBM3 at 2.88 clips/min, 7.98x the slowest card we measured (NVIDIA GeForce RTX 4060 Ti 16GB); best consumer result NVIDIA GeForce RTX 5090 at 1.34; it used about 13.3GB of VRAM.
Which model to pick. On the same card, the NVIDIA H100 80GB HBM3, LTX-Video (image to video) runs at 10.62 clips/min in about 12.4GB; Stable Video Diffusion runs at 4.88 clips/min in about 11.1GB; Wan 2.2 TI2V-5B (image to video) runs at 4.13 clips/min in about 15.3GB; Stable Video Diffusion XT runs at 2.88 clips/min in about 13.3GB; CogVideoX-5B I2V runs at 0.88 clips/min in about 15.9GB; Wan 2.1 I2V 14B (480p) runs at 0.4 clips/min in about 51.9GB; Wan 2.2 I2V A14B runs at 0.4 clips/min in about 74.1GB. LTX-Video (image to video) gets through the work 26.4x as fast as Wan 2.2 I2V A14B, so the model you choose moves the speed as much as the card does. Check the licence column in the fit table before you build on one: Apache-2.0, CogVideoX licence (registration), LTX Community (revenue cap), NVIDIA Open Model, Stability Community (revenue cap) are not the same deal.
About Stable Video Diffusion. Stable Video Diffusion: from stabilityai, 1.5B parameters, on Hugging Face since November 2023. 40,865 downloads in the last 30 days.
How it compares. H100 80GB HBM3: Stable Video Diffusion 4.88 clips/min, Stable Video Diffusion XT 2.88 (2B), Wan 2.2 TI2V-5B (image to video) 4.13 (5B), CogVideoX-5B I2V 0.88 (6B), Wan 2.2 I2V A14B 0.4 (14B). Stable Video Diffusion beats all 4 here.
Cost on a rented GPU. 100 clips of Stable Video Diffusion: $0.14 on a RTX 3060 ($0.036/hr, 3.9 hours), $0.73 on a H100 80GB HBM3 ($2.14/hr, 21 min, 5.2x the cost).
Stable Video Diffusion: cost per 100 clips on rented GPUs
| GPU | Cheapest rate | Speed (clips/min) | Cost per 100 clips |
|---|---|---|---|
| NVIDIA GeForce RTX 3060 | $0.036/hr | 0.43 | $0.14 |
| GeForce RTX 5070 Ti | $0.15/hr | 1.4 | $0.18 |
| NVIDIA GeForce RTX 3090 | $0.12/hr | 1.06 | $0.19 |
| GeForce RTX 5080 | $0.21/hr | 1.67 | $0.21 |
| NVIDIA GeForce RTX 5090 | $0.39/hr | 2.84 | $0.23 |
| NVIDIA GeForce RTX 4080 | $0.20/hr | 1.44 | $0.23 |
| NVIDIA GeForce RTX 4090 | $0.34/hr | 2.23 | $0.25 |
| GeForce RTX 5060 Ti | $0.14/hr | 0.67 | $0.34 |
| NVIDIA L40S | $0.79/hr | 2.19 | $0.60 |
| NVIDIA H100 80GB HBM3 | $2.14/hr | 4.88 | $0.73 |
| NVIDIA L4 | $0.44/hr | 0.63 | $1.17 |
Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.
Speed tiers for Stable Video Diffusion. 4+ clips/min: 1 (H100 80GB HBM3); 1-4 clips/min: 8 (RTX 5090, RTX 4090, RTX 5080); under 1 clips/min: 4 (RTX 5060 Ti, RTX 4060 Ti 16GB, RTX 3060). 4 clips/min is 15 seconds per clip.
VRAM for Stable Video Diffusion. Measured peak 10.6GB, so 12GB is the smallest common card size; smallest card it ran on: RTX 4070 Super (12GB).
Power on Stable Video Diffusion. Most efficient: L4, 72W, 0.19 kWh per 100 clips. Hungriest: H100 80GB HBM3, 615W, 0.21 kWh. At $0.15/kWh: $0.028 per 100 clips.
For image-to-video, the NVIDIA H100 80GB HBM3 is the fastest card we measured. Of the cards you can buy at retail, the NVIDIA GeForce RTX 5090 leads. Check the fit table before buying: VRAM, not speed, is what rules a card out.
Each model runs a fixed workload on every card: a warmup, then timed runs with power, temperature and VRAM sampled every half second through NVML. Models run at the precision and settings from their model card. Datacenter cards run on Modal; consumer cards on rented machines. Non-commercially licensed models are not part of this page.