CogVideoX-2B · 13 GPUs measured first-party · text-to-video · Updated October 2026

What GPU Do You Need for CogVideoX-2B?

CogVideoX-2B on 13 GPUs, measured first-party: NVIDIA H100 80GB HBM3 leads at 1.17 frames/s, L4 trails at 0.179 frames/s, and it peaked at 20GB of VRAM.

Benchmarked weights: THUDM/CogVideoX-2b

Fastest we measured
NVIDIA H100 80GB HBM3

NVIDIA H100 80GB HBM3

1.17 frames/s on CogVideoX-2B, the ceiling. Measured on our bench. 80GB of VRAM, $30,000 at launch.

Pros
  • 1.17 frames/s on CogVideoX-2B
  • 80GB, clears the CogVideoX-2B floor
  • Rentable by the hour rather than bought
Cons
  • 700W board rating
  • Datacenter or workstation hardware, not a retail purchase
Best consumer card
NVIDIA GeForce RTX 5090

NVIDIA GeForce RTX 5090

0.7 frames/s on CogVideoX-2B, fastest card you can buy at retail. Measured on our bench. 32GB of VRAM, $1,999 at launch.

Pros
  • 0.7 frames/s on CogVideoX-2B
  • 32GB, clears the CogVideoX-2B floor
Cons
  • 575W board rating
Cheapest card that runs it
NVIDIA GeForce RTX 3060

NVIDIA GeForce RTX 3060

0.09 frames/s on CogVideoX-2B, lowest launch price that still fits. Measured on our bench. 12GB of VRAM, $329 at launch.

Pros
  • 0.09 frames/s on CogVideoX-2B
  • 12GB, clears the CogVideoX-2B floor
Cons
  • 170W board rating
Best value
GeForce RTX 5070 Ti

GeForce RTX 5070 Ti

0.29 frames/s on CogVideoX-2B, most speed per dollar. Measured on our bench. 16GB of VRAM, $749 at launch. That is 0.39 frames/s per $1,000 of launch price.

Pros
  • 0.29 frames/s on CogVideoX-2B
  • 16GB, clears the CogVideoX-2B floor
Cons
  • 300W board rating
1.17frames/s
Fastest: NVIDIA H100 80GB HBM3
measured, 3-run average
~10GB
VRAM needed (measured peak)
GPU-independent, applies to every card
13
GPUs measured
same pinned harness
0.0frames/s/W
Most efficient: NVIDIA L4
real power sampling, not TDP

What GPU Do You Need for CogVideoX-2B?, frames/s by GPU

NVIDIA H100 80GB HBM3
1.17 frames/s
NVIDIA GeForce RTX 5090
0.7 frames/s
NVIDIA L40S
0.6 frames/s
NVIDIA GeForce RTX 4090
0.52 frames/s
GeForce RTX 5080
0.36 frames/s
NVIDIA GeForce RTX 4080
0.31 frames/s
GeForce RTX 5070 Ti
0.29 frames/s
NVIDIA GeForce RTX 3090
0.24 frames/s
NVIDIA GeForce RTX 4070 Super
0.23 frames/s
NVIDIA L4
0.18 frames/s
GeForce RTX 5060 Ti
0.16 frames/s
NVIDIA GeForce RTX 4060 Ti 16GB
0.14 frames/s
NVIDIA GeForce RTX 3060
0.09 frames/s

Efficiency: frames/s per 100W drawn

NVIDIA L4
0.25 frames/s / 100W
NVIDIA L40S
0.18 frames/s / 100W
NVIDIA H100 80GB HBM3
0.18 frames/s / 100W
GeForce RTX 5070 Ti
0.13 frames/s / 100W
GeForce RTX 5080
0.13 frames/s / 100W
NVIDIA GeForce RTX 4090
0.13 frames/s / 100W
NVIDIA GeForce RTX 4070 Super
0.12 frames/s / 100W
GeForce RTX 5060 Ti
0.12 frames/s / 100W
NVIDIA GeForce RTX 4080
0.11 frames/s / 100W
NVIDIA GeForce RTX 4060 Ti 16GB
0.1 frames/s / 100W
NVIDIA GeForce RTX 3090
0.07 frames/s / 100W
NVIDIA GeForce RTX 3060
0.06 frames/s / 100W

Power is the average pulled during the run, sampled at 1Hz. The fastest card is often not the one here, and for anything left running this is the number that shows up on the bill.

Value: frames/s per $1,000 of MSRP

GeForce RTX 5070 Ti
0.39 frames/s / $1k
NVIDIA GeForce RTX 4070 Super
0.38 frames/s / $1k
GeForce RTX 5060 Ti
0.36 frames/s / $1k
GeForce RTX 5080
0.36 frames/s / $1k
NVIDIA GeForce RTX 5090
0.35 frames/s / $1k
NVIDIA GeForce RTX 4090
0.32 frames/s / $1k
NVIDIA GeForce RTX 4060 Ti 16GB
0.29 frames/s / $1k
NVIDIA GeForce RTX 3060
0.27 frames/s / $1k
NVIDIA GeForce RTX 4080
0.26 frames/s / $1k
NVIDIA GeForce RTX 3090
0.16 frames/s / $1k
NVIDIA L40S
0.08 frames/s / $1k
NVIDIA L4
0.07 frames/s / $1k
NVIDIA H100 80GB HBM3
0.04 frames/s / $1k

Launch price, not street price, so it ages. A speed leaderboard always crowns the most expensive card; this is the counterweight.

CogVideoX-2B. Measured text-to-video speed by GPU

NVIDIA H100 80GB HBM31.17
NVIDIA GeForce RTX 50900.7
NVIDIA L40S0.6
NVIDIA GeForce RTX 40900.52
GeForce RTX 50800.36
NVIDIA GeForce RTX 40800.31
GeForce RTX 5070 Ti0.29
NVIDIA GeForce RTX 30900.24
NVIDIA GeForce RTX 4070 Super0.23
NVIDIA L40.18
GeForce RTX 5060 Ti0.16
NVIDIA GeForce RTX 4060 Ti 16GB0.14
NVIDIA GeForce RTX 30600.09
GPUframes/sPeak VRAMAvg power
NVIDIA H100 80GB HBM31.1734.2GB666.3 W
NVIDIA GeForce RTX 50900.7——
NVIDIA L40S0.634.1GB337.9 W
NVIDIA GeForce RTX 40900.5219.5GB409.6 W
GeForce RTX 50800.369.4GB281.3 W
NVIDIA GeForce RTX 40800.319.3GB294.0 W
GeForce RTX 5070 Ti0.299.3GB220.1 W
NVIDIA GeForce RTX 30900.2422.2GB348.0 W
NVIDIA GeForce RTX 4070 Super0.239.2GB193.3 W
NVIDIA L40.1818.6GB71.6 W
GeForce RTX 5060 Ti0.169.2GB135.2 W
NVIDIA GeForce RTX 4060 Ti 16GB0.149.1GB144.8 W
NVIDIA GeForce RTX 30600.099.1GB155.2 W

What the numbers show. Across 3 GPUs measured on our own bench, H100 80GB HBM3 is fastest at 1.17 frames/s. The slowest, L4, manages 0.18, so the spread is 6.5x from top to bottom. L4 is the most efficient, 0.18 frames/s at 72W. Per dollar of launch price, L40S gives the most (0.1 frames/s per $1,000).

About CogVideoX-2B. CogVideoX-2B: from zai-org, 1.7B parameters, on Hugging Face since August 2024, Apache 2.0 licence. 16,190 downloads in the last 30 days.

How it compares. RTX 4090: CogVideoX-2B 0.52 frames/s, LTX-Video (distilled) 6.7 (2B), Wan 2.1 1.3B 0.67 (1B), CogVideoX-5B 0.17 (6B), Wan 2.2 5B (720p) 0.43. 2 of 4 beat CogVideoX-2B here.

Cost on a rented GPU. 1 minute of 24fps video of CogVideoX-2B: $0.16 on a RTX 3060 ($0.036/hr, 4.5 hours), $0.73 on a H100 80GB HBM3 ($2.14/hr, 21 min, 4.5x the cost).

CogVideoX-2B: cost per 1 minute of 24fps video on rented GPUs

NVIDIA GeForce RTX 3060$0.036/hr
NVIDIA GeForce RTX 3090$0.12/hr
GeForce RTX 5070 Ti$0.15/hr
NVIDIA GeForce RTX 5090$0.39/hr
GeForce RTX 5080$0.21/hr
NVIDIA GeForce RTX 4090$0.34/hr
NVIDIA GeForce RTX 4080$0.20/hr
GeForce RTX 5060 Ti$0.14/hr
NVIDIA L40S$0.79/hr
NVIDIA H100 80GB HBM3$2.14/hr
NVIDIA L4$0.44/hr
GPUCheapest rateSpeed (frames/s)Cost per 1 minute of 24fps video
NVIDIA GeForce RTX 3060$0.036/hr0.09$0.16
NVIDIA GeForce RTX 3090$0.12/hr0.24$0.20
GeForce RTX 5070 Ti$0.15/hr0.29$0.20
NVIDIA GeForce RTX 5090$0.39/hr0.7$0.22
GeForce RTX 5080$0.21/hr0.36$0.23
NVIDIA GeForce RTX 4090$0.34/hr0.52$0.26
NVIDIA GeForce RTX 4080$0.20/hr0.31$0.26
GeForce RTX 5060 Ti$0.14/hr0.16$0.35
NVIDIA L40S$0.79/hr0.6$0.53
NVIDIA H100 80GB HBM3$2.14/hr1.17$0.73
NVIDIA L4$0.44/hr0.18$0.98

Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.

Speed tiers for CogVideoX-2B. 1-10 frames/s: 1 (H100 80GB HBM3); under 1 frames/s: 12 (RTX 5090, RTX 4090, RTX 5080). 10 frames/s turns out a minute of 24fps video in under 2.5 minutes.

VRAM for CogVideoX-2B. Measured peak 9.1GB, so 12GB is the smallest common card size; smallest card it ran on: RTX 4070 Super (12GB).

Power on CogVideoX-2B. Most efficient: L4, 72W, 0.16 kWh per 1 minute of 24fps video. Hungriest: H100 80GB HBM3, 666W, 0.23 kWh. At $0.15/kWh: $0.024 per 1 minute of 24fps video.

Our verdict

Fastest on CogVideoX-2B: NVIDIA H100 80GB HBM3, 1.17 frames/s. Best desktop card: NVIDIA GeForce RTX 5090, 0.7 frames/s. Cheapest consumer card that ran it: NVIDIA GeForce RTX 3060 ($329, 0.09 frames/s). Cheapest to rent per job: NVIDIA GeForce RTX 3060, $0.16 per 1 minute of 24fps video.

FAQ

What GPU do I need to run CogVideoX-2B?
About 10GB. Cheapest consumer card that ran it: NVIDIA GeForce RTX 3060 (12GB, 0.09 frames/s).
How fast is CogVideoX-2B on the NVIDIA GeForce RTX 5090?
0.7 frames/s, 60% of the NVIDIA H100 80GB HBM3.
How much does it cost to run CogVideoX-2B in the cloud?
$0.16 per 1 minute of 24fps video on a NVIDIA GeForce RTX 3060 at $0.036/hr, cheapest of 11 rentable cards we measured.
Can I run CogVideoX-2B on a 12GB, 16GB or 24GB card?
It used 9.1GB at the precision we tested. 12GB: yes; 16GB: yes; 24GB: yes.
Is the RTX 4090 or the RTX 3090 faster for CogVideoX-2B?
The RTX 4090: 0.52 vs 0.24 frames/s, 115% faster on our bench.
Should I buy or rent a RTX 5090 for CogVideoX-2B?
Its $1,999 launch price buys 5,139 rented hours at $0.39/hr, enough for about 8,993x 1 minute of 24fps video of CogVideoX-2B. Buy only if you'll run more than that.

How we test

CogVideoX-2B in diffusers, bf16, one warmup clip and three timed clips from the same prompt, measured as frames of finished video per second, with CPU offload only on cards below the model's full-GPU size. Text-to-video is compute-bound and memory-hungry; cards that need CPU offload lose far more speed than their raw compute suggests.