AI background removal · 1 model · 11 GPUs measured first-party · Updated October 2026
Which graphics card to use for AI background removal, from first-party measurements of BiRefNet on 11 GPUs.

1457.7 images/min on BiRefNet, the ceiling. Measured on our bench. 141GB of VRAM, $31,000 at launch.

1445.9 images/min on BiRefNet, lowest launch price that still fits. Measured on our bench. 96GB of VRAM, $8,565 at launch.
Background removal cuts the subject out of a photo with a clean edge: product shots, profile pictures, thumbnails. BiRefNet is one of the best open models for it, and it is a classic batch job where throughput per dollar matters more than the fastest single image.
We measured 1 model for AI background removal on 11 GPUs. Speed is images processed per minute. Every number below is a first-party run on our own harness; cards absent from a model's chart have not been run on it yet.
BiRefNet: images/min by GPU
Which models fit which card, for AI background removal
| Model | VRAM used | 8GB card | 12GB card | 16GB card | 24GB card | 32GB card | Licence |
|---|---|---|---|---|---|---|---|
| BiRefNet | 3.2GB | Yes | Yes | Yes | Yes | Yes | MIT |
From the lowest VRAM peak we measured for each model, plus 5% headroom. 'No' means it did not fit in that much memory at our settings, not that no setting ever could.
BiRefNet: every GPU we measured
| GPU | images/min | VRAM | Power |
|---|---|---|---|
| NVIDIA H200 | 1457.7 | 141GB | 134.4 W |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 1445.9 | 96GB | 100.3 W |
| NVIDIA H100 80GB HBM3 | 1311 | 80GB | 122.7 W |
| NVIDIA B200 | 1082 | 192GB | 252.1 W |
| NVIDIA B300 | 1064.7 | 288GB | 239.8 W |
| NVIDIA A100 80GB SXM4 | 916.7 | 80GB | 83.2 W |
| NVIDIA L40S | 863.8 | 48GB | 92.1 W |
| NVIDIA A100 40GB SXM4 | 611.6 | 40GB | 75.5 W |
| NVIDIA A10G | 432.3 | 24GB | 68.3 W |
| NVIDIA L4 | 305.8 | 24GB | 45.5 W |
| NVIDIA T4 | 183.5 | 16GB | 52.3 W |
What the numbers show.
BiRefNet: fastest on the NVIDIA H200 at 1457.7 images/min, 7.94x the slowest card we measured (NVIDIA T4); it used about 3.2GB of VRAM.
How it compares. L40S: BiRefNet 863.8 images/min, Depth Anything V2 Large 959.3, Depth Anything V2 Small 970.6, SAM ViT-Base 998.8, SDXL Turbo 423.5 (3B). 3 of 4 beat BiRefNet here.
Cost on a rented GPU. 1,000 images of BiRefNet: $0.012 on a T4 ($0.14/hr, 5 min), $0.041 on a H200 ($3.59/hr, 1 min, 3.3x the cost).
BiRefNet: cost per 1,000 images on rented GPUs
| GPU | Cheapest rate | Speed (images/min) | Cost per 1,000 images |
|---|---|---|---|
| NVIDIA T4 | $0.14/hr | 183.5 | $0.012 |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | $1.08/hr | 1445.9 | $0.012 |
| NVIDIA A100 40GB SXM4 | $0.47/hr | 611.6 | $0.013 |
| NVIDIA L40S | $0.79/hr | 863.8 | $0.015 |
| NVIDIA A100 80GB SXM4 | $0.95/hr | 916.7 | $0.017 |
| NVIDIA L4 | $0.44/hr | 305.8 | $0.024 |
| NVIDIA H100 80GB HBM3 | $2.14/hr | 1311 | $0.027 |
| NVIDIA H200 | $3.59/hr | 1457.7 | $0.041 |
| NVIDIA B200 | $5.98/hr | 1082 | $0.092 |
| NVIDIA B300 | $6.94/hr | 1064.7 | $0.11 |
Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.
Speed tiers for BiRefNet. 30+ images/min: 11 (H200, RTX PRO 6000 Blackwell Workstation Edition, H100 80GB HBM3). 30 images/min means two seconds or less per picture.
Time per image. BiRefNet: 0.0s per image on the H200, 0.3s on the T4. A batch of 100 takes 0 min on the fastest card and 1 min on the slowest.
VRAM for BiRefNet. Measured peak 3.0GB, so 8GB is the smallest common card size; smallest card it ran on: T4 (16GB).
Power on BiRefNet. Most efficient: A100 80GB SXM4, 83W, 1.5 Wh per 1,000 images. Hungriest: B200, 252W, 3.9 Wh.
For AI background removal, the NVIDIA H200 is the fastest card we measured. We have not measured a consumer card on this job yet; the picks above are datacenter and workstation hardware. Check the fit table before buying: VRAM, not speed, is what rules a card out.
Each model runs a fixed workload on every card: a warmup, then timed runs with power, temperature and VRAM sampled every half second through NVML. Models run at the precision and settings from their model card. Datacenter cards run on Modal; consumer cards on rented machines. Non-commercially licensed models are not part of this page.