Wan 2.1 1.3B · 12 GPUs measured first-party · text-to-video · Updated October 2026

What GPU Do You Need for Wan 2.1 1.3B?

Wan 2.1 1.3B on 12 GPUs, measured first-party: NVIDIA L40S leads at 0.73 frames/s, RTX 3060 trails at 0.122 frames/s, and it peaked at 12GB of VRAM.

Benchmarked weights: Wan-AI/Wan2.1-T2V-1.3B-Diffusers

Fastest we measured
NVIDIA GeForce RTX 5090

NVIDIA GeForce RTX 5090

0.88 frames/s on Wan 2.1 1.3B, the ceiling. Measured on our bench. 32GB of VRAM, $1,999 at launch.

Pros
  • 0.88 frames/s on Wan 2.1 1.3B
  • 32GB, clears the Wan 2.1 1.3B floor
Cons
  • 575W board rating
Best consumer card
NVIDIA GeForce RTX 4090

NVIDIA GeForce RTX 4090

0.67 frames/s on Wan 2.1 1.3B, fastest card you can buy at retail. Measured on our bench. 24GB of VRAM, $1,599 at launch.

Pros
  • 0.67 frames/s on Wan 2.1 1.3B
  • 24GB, clears the Wan 2.1 1.3B floor
Cons
  • 450W board rating
Cheapest card that runs it
NVIDIA GeForce RTX 3060

NVIDIA GeForce RTX 3060

0.12 frames/s on Wan 2.1 1.3B, lowest launch price that still fits. Measured on our bench. 12GB of VRAM, $329 at launch.

Pros
  • 0.12 frames/s on Wan 2.1 1.3B
  • 12GB, clears the Wan 2.1 1.3B floor
Cons
  • 170W board rating
Best value
GeForce RTX 5070 Ti

GeForce RTX 5070 Ti

0.37 frames/s on Wan 2.1 1.3B, most speed per dollar. Measured on our bench. 16GB of VRAM, $749 at launch. That is 0.5 frames/s per $1,000 of launch price.

Pros
  • 0.37 frames/s on Wan 2.1 1.3B
  • 16GB, clears the Wan 2.1 1.3B floor
Cons
  • 300W board rating
0.88frames/s
Fastest: NVIDIA GeForce RTX 5090
measured, 3-run average
~12GB
VRAM needed (measured peak)
GPU-independent, applies to every card
12
GPUs measured
same pinned harness
0.0frames/s/W
Most efficient: NVIDIA L4
real power sampling, not TDP

What GPU Do You Need for Wan 2.1 1.3B?, frames/s by GPU

NVIDIA GeForce RTX 5090
0.88 frames/s
NVIDIA L40S
0.73 frames/s
NVIDIA GeForce RTX 4090
0.67 frames/s
GeForce RTX 5080
0.45 frames/s
NVIDIA GeForce RTX 4080
0.38 frames/s
GeForce RTX 5070 Ti
0.37 frames/s
NVIDIA GeForce RTX 3090
0.34 frames/s
NVIDIA GeForce RTX 4070 Super
0.28 frames/s
NVIDIA L4
0.22 frames/s
GeForce RTX 5060 Ti
0.2 frames/s
NVIDIA GeForce RTX 4060 Ti 16GB
0.18 frames/s
NVIDIA GeForce RTX 3060
0.12 frames/s

Efficiency: frames/s per 100W drawn

NVIDIA L4
0.31 frames/s / 100W
NVIDIA L40S
0.21 frames/s / 100W
GeForce RTX 5070 Ti
0.19 frames/s / 100W
GeForce RTX 5080
0.18 frames/s / 100W
GeForce RTX 5060 Ti
0.16 frames/s / 100W
NVIDIA GeForce RTX 4090
0.16 frames/s / 100W
NVIDIA GeForce RTX 4070 Super
0.16 frames/s / 100W
NVIDIA GeForce RTX 5090
0.16 frames/s / 100W
NVIDIA GeForce RTX 4080
0.14 frames/s / 100W
NVIDIA GeForce RTX 4060 Ti 16GB
0.13 frames/s / 100W
NVIDIA GeForce RTX 3090
0.1 frames/s / 100W
NVIDIA GeForce RTX 3060
0.08 frames/s / 100W

Power is the average pulled during the run, sampled at 1Hz. The fastest card is often not the one here, and for anything left running this is the number that shows up on the bill.

Value: frames/s per $1,000 of MSRP

GeForce RTX 5070 Ti
0.5 frames/s / $1k
NVIDIA GeForce RTX 4070 Super
0.47 frames/s / $1k
GeForce RTX 5060 Ti
0.46 frames/s / $1k
GeForce RTX 5080
0.45 frames/s / $1k
NVIDIA GeForce RTX 5090
0.44 frames/s / $1k
NVIDIA GeForce RTX 4090
0.42 frames/s / $1k
NVIDIA GeForce RTX 3060
0.37 frames/s / $1k
NVIDIA GeForce RTX 4060 Ti 16GB
0.35 frames/s / $1k
NVIDIA GeForce RTX 4080
0.32 frames/s / $1k
NVIDIA GeForce RTX 3090
0.23 frames/s / $1k
NVIDIA L40S
0.1 frames/s / $1k
NVIDIA L4
0.09 frames/s / $1k

Launch price, not street price, so it ages. A speed leaderboard always crowns the most expensive card; this is the counterweight.

Wan 2.1 1.3B. Measured text-to-video speed by GPU

NVIDIA GeForce RTX 50900.88
NVIDIA L40S0.73
NVIDIA GeForce RTX 40900.67
GeForce RTX 50800.45
NVIDIA GeForce RTX 40800.38
GeForce RTX 5070 Ti0.37
NVIDIA GeForce RTX 30900.34
NVIDIA GeForce RTX 4070 Super0.28
NVIDIA L40.22
GeForce RTX 5060 Ti0.2
NVIDIA GeForce RTX 4060 Ti 16GB0.18
NVIDIA GeForce RTX 30600.12
GPUframes/sPeak VRAMAvg power
NVIDIA GeForce RTX 50900.8821.0GB566.1 W
NVIDIA L40S0.7319.9GB346.3 W
NVIDIA GeForce RTX 40900.6720.1GB411.2 W
GeForce RTX 50800.4511.3GB253.2 W
NVIDIA GeForce RTX 40800.3811.2GB271.1 W
GeForce RTX 5070 Ti0.3711.2GB195.1 W
NVIDIA GeForce RTX 30900.3420.6GB348.3 W
NVIDIA GeForce RTX 4070 Super0.2811.2GB179.2 W
NVIDIA L40.2219.5GB71.7 W
GeForce RTX 5060 Ti0.211.1GB120.9 W
NVIDIA GeForce RTX 4060 Ti 16GB0.1811.1GB139.3 W
NVIDIA GeForce RTX 30600.1211.1GB144.6 W

What the numbers show. Across 5 GPUs measured on our own bench, L40S is fastest at 0.73 frames/s. The slowest, RTX 3060, manages 0.12, so the spread is 6.0x from top to bottom. L4 is the most efficient, 0.22 frames/s at 72W. Per dollar of launch price, RTX 4090 gives the most (0.4 frames/s per $1,000). The fastest card with 16GB or less is RTX 4060 Ti 16GB at 0.18 frames/s.

About Wan 2.1 1.3B. Wan 2.1 1.3B: from Wan-AI, 1.4B parameters, on Hugging Face since February 2025, Apache 2.0 licence. 314,837 downloads in the last 30 days and 7 community quantizations.

How it compares. RTX 4090: Wan 2.1 1.3B 0.67 frames/s, CogVideoX-2B 0.52 (2B), LTX-Video (distilled) 6.7 (2B), CogVideoX-5B 0.17 (6B), Wan 2.2 5B (720p) 0.43. 1 of 4 beat Wan 2.1 1.3B here.

Cost on a rented GPU. 1 minute of 24fps video of Wan 2.1 1.3B: $0.12 on a RTX 3060 ($0.036/hr, 3.3 hours), $0.18 on a RTX 5090 ($0.39/hr, 27 min, 1.5x the cost).

Wan 2.1 1.3B: cost per 1 minute of 24fps video on rented GPUs

NVIDIA GeForce RTX 3060$0.036/hr
NVIDIA GeForce RTX 3090$0.12/hr
GeForce RTX 5070 Ti$0.15/hr
NVIDIA GeForce RTX 5090$0.39/hr
GeForce RTX 5080$0.21/hr
NVIDIA GeForce RTX 4090$0.34/hr
NVIDIA GeForce RTX 4080$0.20/hr
GeForce RTX 5060 Ti$0.14/hr
NVIDIA L40S$0.79/hr
NVIDIA L4$0.44/hr
GPUCheapest rateSpeed (frames/s)Cost per 1 minute of 24fps video
NVIDIA GeForce RTX 3060$0.036/hr0.12$0.12
NVIDIA GeForce RTX 3090$0.12/hr0.34$0.14
GeForce RTX 5070 Ti$0.15/hr0.37$0.16
NVIDIA GeForce RTX 5090$0.39/hr0.88$0.18
GeForce RTX 5080$0.21/hr0.45$0.19
NVIDIA GeForce RTX 4090$0.34/hr0.67$0.20
NVIDIA GeForce RTX 4080$0.20/hr0.38$0.21
GeForce RTX 5060 Ti$0.14/hr0.2$0.27
NVIDIA L40S$0.79/hr0.73$0.43
NVIDIA L4$0.44/hr0.22$0.80

Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.

Speed tiers for Wan 2.1 1.3B. under 1 frames/s: 12 (RTX 5090, RTX 4090, RTX 5080). 10 frames/s turns out a minute of 24fps video in under 2.5 minutes.

VRAM for Wan 2.1 1.3B. Measured peak 11.1GB, so 12GB is the smallest common card size; smallest card it ran on: RTX 4070 Super (12GB).

Power on Wan 2.1 1.3B. Most efficient: L4, 72W, 0.13 kWh per 1 minute of 24fps video. Hungriest: RTX 5090, 566W, 0.26 kWh. At $0.15/kWh: $0.020 per 1 minute of 24fps video.

Our verdict

Fastest on Wan 2.1 1.3B: NVIDIA GeForce RTX 5090, 0.88 frames/s. Cheapest consumer card that ran it: NVIDIA GeForce RTX 3060 ($329, 0.12 frames/s). Cheapest to rent per job: NVIDIA GeForce RTX 3060, $0.12 per 1 minute of 24fps video.

FAQ

What GPU do I need to run Wan 2.1 1.3B?
About 12GB. Cheapest consumer card that ran it: NVIDIA GeForce RTX 3060 (12GB, 0.12 frames/s).
How fast is Wan 2.1 1.3B on the NVIDIA GeForce RTX 5090?
0.88 frames/s at 566W, faster than every datacenter card we ran it on.
How much does it cost to run Wan 2.1 1.3B in the cloud?
$0.12 per 1 minute of 24fps video on a NVIDIA GeForce RTX 3060 at $0.036/hr, cheapest of 10 rentable cards we measured.
Can I run Wan 2.1 1.3B on a 12GB, 16GB or 24GB card?
It used 11.1GB at the precision we tested. 12GB: yes; 16GB: yes; 24GB: yes.
Is the RTX 4090 or the RTX 3090 faster for Wan 2.1 1.3B?
The RTX 4090: 0.67 vs 0.34 frames/s, 97% faster on our bench.
Should I buy or rent a RTX 5090 for Wan 2.1 1.3B?
Its $1,999 launch price buys 5,139 rented hours at $0.39/hr, enough for about 11,293x 1 minute of 24fps video of Wan 2.1 1.3B. Buy only if you'll run more than that.

How we test

Wan 2.1 1.3B in diffusers, bf16, one warmup clip and three timed clips from the same prompt, measured as frames of finished video per second, with CPU offload only on cards below the model's full-GPU size. Text-to-video is compute-bound and memory-hungry; cards that need CPU offload lose far more speed than their raw compute suggests.