Vision-language models · 2 models · 6 GPUs measured first-party · Updated October 2026
Which graphics card to use for vision-language models, from first-party measurements of Florence-2 Base, Florence-2 Large on 6 GPUs.

249.6 images/min on Florence-2 Base, the ceiling. Measured on our bench. 80GB of VRAM, $30,000 at launch.
Vision-language models read images: captions, OCR, object detection and visual question answering from one model. Microsoft's Florence-2 is a compact open family widely used for captioning and labeling datasets.
We measured 2 models for vision-language models on 6 GPUs. Speed is images processed per minute. Every number below is a first-party run on our own harness; cards absent from a model's chart have not been run on it yet.
Florence-2 Base: images/min by GPU
Florence-2 Large: images/min by GPU
Which models fit which card, for vision-language models
| Model | VRAM used | 8GB card | 12GB card | 16GB card | 24GB card | 32GB card | Licence |
|---|---|---|---|---|---|---|---|
| Florence-2 Base | 1.3GB | Yes | Yes | Yes | Yes | Yes | MIT |
| Florence-2 Large | 2.5GB | Yes | Yes | Yes | Yes | Yes | MIT |
From the lowest VRAM peak we measured for each model, plus 5% headroom. 'No' means it did not fit in that much memory at our settings, not that no setting ever could.
Every GPU x every vision-language models model (images/min)
| GPU | Florence-2 Base | Florence-2 Large |
|---|---|---|
| NVIDIA H100 80GB HBM3 | 249.6 | 156.2 |
| NVIDIA A10G | 210.8 | 117.9 |
| NVIDIA L40S | 184.2 | 106.5 |
| NVIDIA L4 | 174.7 | 93.77 |
| NVIDIA T4 | 148.6 | 83.43 |
| NVIDIA A100 80GB SXM4 | 126.4 | 70.94 |
— = not measured on that card yet.
What the numbers show.
Florence-2 Base: fastest on the NVIDIA H100 80GB HBM3 at 249.6 images/min, 1.97x the slowest card we measured (NVIDIA A100 80GB SXM4); it used about 1.3GB of VRAM.
Florence-2 Large: fastest on the NVIDIA H100 80GB HBM3 at 156.2 images/min, 2.2x the slowest card we measured (NVIDIA A100 80GB SXM4); it used about 2.5GB of VRAM.
Which model to pick. On the same card, the NVIDIA H100 80GB HBM3, Florence-2 Base runs at 249.6 images/min in about 1.3GB; Florence-2 Large runs at 156.2 images/min in about 2.5GB. Florence-2 Base gets through the work 1.6x as fast as Florence-2 Large, so the model you choose moves the speed as much as the card does.
How it compares. L40S: Florence-2 Base 184.2 images/min, SAM ViT-Huge 238.1, Florence-2 Large 106.5, Stable Diffusion 1.5 58.54, Sana 1.6B 50.81. 1 of 4 beat Florence-2 Base here.
Cost on a rented GPU. 1,000 images of Florence-2 Base: $0.015 on a T4 ($0.14/hr, 7 min), $0.14 on a H100 80GB HBM3 ($2.14/hr, 4 min, 9.4x the cost).
Florence-2 Base: cost per 1,000 images on rented GPUs
| GPU | Cheapest rate | Speed (images/min) | Cost per 1,000 images |
|---|---|---|---|
| NVIDIA T4 | $0.14/hr | 148.6 | $0.015 |
| NVIDIA L4 | $0.44/hr | 174.7 | $0.042 |
| NVIDIA L40S | $0.79/hr | 184.2 | $0.071 |
| NVIDIA A100 80GB SXM4 | $0.95/hr | 126.4 | $0.12 |
| NVIDIA H100 80GB HBM3 | $2.14/hr | 249.6 | $0.14 |
Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.
Speed tiers for Florence-2 Base. 30+ images/min: 6 (H100 80GB HBM3, A10G, L40S). 30 images/min means two seconds or less per picture.
Time per image. Florence-2 Base: 0.2s per image on the H100 80GB HBM3, 0.5s on the A100 80GB SXM4. A batch of 100 takes 0 min on the fastest card and 1 min on the slowest.
VRAM for Florence-2 Base. Measured peak 1.2GB, so 8GB is the smallest common card size; smallest card it ran on: T4 (16GB).
Power on Florence-2 Base. Most efficient: L4, 38W, 3.7 Wh per 1,000 images. Hungriest: L40S, 99W, 9.0 Wh.
For vision-language models, the NVIDIA H100 80GB HBM3 is the fastest card we measured. We have not measured a consumer card on this job yet; the picks above are datacenter and workstation hardware. Check the fit table before buying: VRAM, not speed, is what rules a card out.
Each model runs a fixed workload on every card: a warmup, then timed runs with power, temperature and VRAM sampled every half second through NVML. Models run at the precision and settings from their model card. Datacenter cards run on Modal; consumer cards on rented machines. Non-commercially licensed models are not part of this page.