Vision-language models · 2 models · 6 GPUs measured first-party · Updated October 2026

Best GPU for Vision-language models

Which graphics card to use for vision-language models, from first-party measurements of Florence-2 Base, Florence-2 Large on 6 GPUs.

Fastest we measured
NVIDIA H100 80GB HBM3

NVIDIA H100 80GB HBM3

249.6 images/min on Florence-2 Base, the ceiling. Measured on our bench. 80GB of VRAM, $30,000 at launch.

Pros
  • 249.6 images/min on Florence-2 Base
  • 80GB, clears the Florence-2 Base floor
  • Rentable by the hour rather than bought
Cons
  • 700W board rating
  • Datacenter or workstation hardware, not a retail purchase
2
Models measured
Florence-2 Base, Florence-2 Large
6
GPUs measured
first-party runs, not spec-sheet estimates
249.6images/min
Fastest: NVIDIA H100 80GB HBM3
on Florence-2 Base
1.3GB
Lightest model's VRAM need
measured peak, +5% headroom

Vision-language models read images: captions, OCR, object detection and visual question answering from one model. Microsoft's Florence-2 is a compact open family widely used for captioning and labeling datasets.

We measured 2 models for vision-language models on 6 GPUs. Speed is images processed per minute. Every number below is a first-party run on our own harness; cards absent from a model's chart have not been run on it yet.

Florence-2 Base: images/min by GPU

NVIDIA H100 80GB HBM3
249.6 images/min
NVIDIA A10G
210.8 images/min
NVIDIA L40S
184.2 images/min
NVIDIA L4
174.7 images/min
NVIDIA T4
148.6 images/min
NVIDIA A100 80GB SXM4
126.4 images/min

Florence-2 Large: images/min by GPU

NVIDIA H100 80GB HBM3
156.2 images/min
NVIDIA A10G
117.9 images/min
NVIDIA L40S
106.5 images/min
NVIDIA L4
93.77 images/min
NVIDIA T4
83.43 images/min
NVIDIA A100 80GB SXM4
70.94 images/min

Which models fit which card, for vision-language models

ModelVRAM used8GB card12GB card16GB card24GB card32GB cardLicence
Florence-2 Base1.3GBYesYesYesYesYesMIT
Florence-2 Large2.5GBYesYesYesYesYesMIT

From the lowest VRAM peak we measured for each model, plus 5% headroom. 'No' means it did not fit in that much memory at our settings, not that no setting ever could.

Every GPU x every vision-language models model (images/min)

NVIDIA H100 80GB HBM3249.6
NVIDIA A10G210.8
NVIDIA L40S184.2
NVIDIA L4174.7
NVIDIA T4148.6
NVIDIA A100 80GB SXM4126.4
GPUFlorence-2 BaseFlorence-2 Large
NVIDIA H100 80GB HBM3249.6156.2
NVIDIA A10G210.8117.9
NVIDIA L40S184.2106.5
NVIDIA L4174.793.77
NVIDIA T4148.683.43
NVIDIA A100 80GB SXM4126.470.94

— = not measured on that card yet.

What the numbers show.

Florence-2 Base: fastest on the NVIDIA H100 80GB HBM3 at 249.6 images/min, 1.97x the slowest card we measured (NVIDIA A100 80GB SXM4); it used about 1.3GB of VRAM.

Florence-2 Large: fastest on the NVIDIA H100 80GB HBM3 at 156.2 images/min, 2.2x the slowest card we measured (NVIDIA A100 80GB SXM4); it used about 2.5GB of VRAM.

Which model to pick. On the same card, the NVIDIA H100 80GB HBM3, Florence-2 Base runs at 249.6 images/min in about 1.3GB; Florence-2 Large runs at 156.2 images/min in about 2.5GB. Florence-2 Base gets through the work 1.6x as fast as Florence-2 Large, so the model you choose moves the speed as much as the card does.

How it compares. L40S: Florence-2 Base 184.2 images/min, SAM ViT-Huge 238.1, Florence-2 Large 106.5, Stable Diffusion 1.5 58.54, Sana 1.6B 50.81. 1 of 4 beat Florence-2 Base here.

Cost on a rented GPU. 1,000 images of Florence-2 Base: $0.015 on a T4 ($0.14/hr, 7 min), $0.14 on a H100 80GB HBM3 ($2.14/hr, 4 min, 9.4x the cost).

Florence-2 Base: cost per 1,000 images on rented GPUs

NVIDIA T4$0.14/hr
NVIDIA L4$0.44/hr
NVIDIA L40S$0.79/hr
NVIDIA A100 80GB SXM4$0.95/hr
NVIDIA H100 80GB HBM3$2.14/hr
GPUCheapest rateSpeed (images/min)Cost per 1,000 images
NVIDIA T4$0.14/hr148.6$0.015
NVIDIA L4$0.44/hr174.7$0.042
NVIDIA L40S$0.79/hr184.2$0.071
NVIDIA A100 80GB SXM4$0.95/hr126.4$0.12
NVIDIA H100 80GB HBM3$2.14/hr249.6$0.14

Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.

Speed tiers for Florence-2 Base. 30+ images/min: 6 (H100 80GB HBM3, A10G, L40S). 30 images/min means two seconds or less per picture.

Time per image. Florence-2 Base: 0.2s per image on the H100 80GB HBM3, 0.5s on the A100 80GB SXM4. A batch of 100 takes 0 min on the fastest card and 1 min on the slowest.

VRAM for Florence-2 Base. Measured peak 1.2GB, so 8GB is the smallest common card size; smallest card it ran on: T4 (16GB).

Power on Florence-2 Base. Most efficient: L4, 38W, 3.7 Wh per 1,000 images. Hungriest: L40S, 99W, 9.0 Wh.

Our verdict

For vision-language models, the NVIDIA H100 80GB HBM3 is the fastest card we measured. We have not measured a consumer card on this job yet; the picks above are datacenter and workstation hardware. Check the fit table before buying: VRAM, not speed, is what rules a card out.

FAQ

What is the fastest GPU for vision-language models?
In our runs, the NVIDIA H100 80GB HBM3 at 249.6 images/min on Florence-2 Base. We measured 2 models on 6 GPUs for this page.
How much VRAM do I need for vision-language models?
The lightest model here, Florence-2 Base, used about 1.3GB. The table above shows which models fit 8, 12, 16, 24 and 32GB cards, from measured peaks.
Are these numbers measured or estimated?
Measured. Every number on this page is a first-party run on our own harness, with power and VRAM sampled during the run. Cards we have not run yet are simply absent, not filled in.
What GPU do I need to run Florence-2 Base?
About 1GB. Smallest card that ran it: NVIDIA T4 (16GB).
How much does it cost to run Florence-2 Base in the cloud?
$0.015 per 1,000 images on a NVIDIA T4 at $0.14/hr, cheapest of 5 rentable cards we measured.
Can I run Florence-2 Base on a 12GB, 16GB or 24GB card?
It used 1.2GB at the precision we tested. 12GB: yes; 16GB: yes; 24GB: yes.
Is the H100 80GB HBM3 or the A100 80GB SXM4 faster for Florence-2 Base?
The H100 80GB HBM3: 249.6 vs 126.4 images/min, 97% faster on our bench.

How we test

Each model runs a fixed workload on every card: a warmup, then timed runs with power, temperature and VRAM sampled every half second through NVML. Models run at the precision and settings from their model card. Datacenter cards run on Modal; consumer cards on rented machines. Non-commercially licensed models are not part of this page.