Depth estimation · 1 model · 11 GPUs measured first-party · Updated October 2026
Which graphics card to use for depth estimation, from first-party measurements of Depth Anything V2 Small on 11 GPUs.

1491.6 images/min on Depth Anything V2 Small, the ceiling. Measured on our bench. 288GB of VRAM, $40,000 at launch.

1374.6 images/min on Depth Anything V2 Small, lowest launch price that still fits. Measured on our bench. 96GB of VRAM, $8,565 at launch.
Depth estimation predicts how far away every pixel in a photo is, from a single image. It feeds 3D effects, video relighting, robotics and ControlNet-guided image generation.
We measured 1 model for depth estimation on 11 GPUs. Speed is images processed per minute. Every number below is a first-party run on our own harness; cards absent from a model's chart have not been run on it yet.
Depth Anything V2 Small: images/min by GPU
Which models fit which card, for depth estimation
| Model | VRAM used | 8GB card | 12GB card | 16GB card | 24GB card | 32GB card | Licence |
|---|---|---|---|---|---|---|---|
| Depth Anything V2 Small | 3.2GB | Yes | Yes | Yes | Yes | Yes | Apache-2.0 |
From the lowest VRAM peak we measured for each model, plus 5% headroom. 'No' means it did not fit in that much memory at our settings, not that no setting ever could.
Depth Anything V2 Small: every GPU we measured
| GPU | images/min | VRAM | Power |
|---|---|---|---|
| NVIDIA B300 | 1491.6 | 288GB | 235.0 W |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 1374.6 | 96GB | 87.5 W |
| NVIDIA H200 | 1198.6 | 141GB | 125.1 W |
| NVIDIA B200 | 1137.5 | 192GB | 190.2 W |
| NVIDIA H100 80GB HBM3 | 1081.5 | 80GB | 115.6 W |
| NVIDIA L40S | 970.6 | 48GB | 85.0 W |
| NVIDIA A10G | 941.2 | 24GB | 55.1 W |
| NVIDIA A100 80GB SXM4 | 925.1 | 80GB | 80.5 W |
| NVIDIA L4 | 916.7 | 24GB | 28.4 W |
| NVIDIA A100 40GB SXM4 | 613.4 | 40GB | 65.6 W |
| NVIDIA T4 | 612.2 | 16GB | 27.4 W |
What the numbers show.
Depth Anything V2 Small: fastest on the NVIDIA B300 at 1491.6 images/min, 2.44x the slowest card we measured (NVIDIA T4); it used about 3.2GB of VRAM.
How it compares. L40S: Depth Anything V2 Small 970.6 images/min, Depth Anything V2 Large 959.3, SAM ViT-Base 998.8, BiRefNet 863.8, SDXL Turbo 423.5 (3B). 1 of 4 beat Depth Anything V2 Small here.
Cost on a rented GPU. 1,000 images of Depth Anything V2 Small: $0.004 on a T4 ($0.14/hr, 2 min), $0.078 on a B300 ($6.94/hr, 1 min, 20.9x the cost).
Depth Anything V2 Small: cost per 1,000 images on rented GPUs
| GPU | Cheapest rate | Speed (images/min) | Cost per 1,000 images |
|---|---|---|---|
| NVIDIA T4 | $0.14/hr | 612.2 | $0.004 |
| NVIDIA L4 | $0.44/hr | 916.7 | $0.008 |
| NVIDIA A100 40GB SXM4 | $0.47/hr | 613.4 | $0.013 |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | $1.08/hr | 1374.6 | $0.013 |
| NVIDIA L40S | $0.79/hr | 970.6 | $0.014 |
| NVIDIA A100 80GB SXM4 | $0.95/hr | 925.1 | $0.017 |
| NVIDIA H100 80GB HBM3 | $2.14/hr | 1081.5 | $0.033 |
| NVIDIA H200 | $3.59/hr | 1198.6 | $0.050 |
| NVIDIA B300 | $6.94/hr | 1491.6 | $0.078 |
| NVIDIA B200 | $5.98/hr | 1137.5 | $0.088 |
Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.
Speed tiers for Depth Anything V2 Small. 30+ images/min: 11 (B300, RTX PRO 6000 Blackwell Workstation Edition, H200). 30 images/min means two seconds or less per picture.
Time per image. Depth Anything V2 Small: 0.0s per image on the B300, 0.1s on the T4. A batch of 100 takes 0 min on the fastest card and 0 min on the slowest.
VRAM for Depth Anything V2 Small. Measured peak 3.0GB, so 8GB is the smallest common card size; smallest card it ran on: T4 (16GB).
Power on Depth Anything V2 Small. Most efficient: L4, 28W, 0.5 Wh per 1,000 images. Hungriest: L40S, 85W, 1.5 Wh.
For depth estimation, the NVIDIA B300 is the fastest card we measured. We have not measured a consumer card on this job yet; the picks above are datacenter and workstation hardware. Check the fit table before buying: VRAM, not speed, is what rules a card out.
Each model runs a fixed workload on every card: a warmup, then timed runs with power, temperature and VRAM sampled every half second through NVML. Models run at the precision and settings from their model card. Datacenter cards run on Modal; consumer cards on rented machines. Non-commercially licensed models are not part of this page.