GPU Rankings By AI Task

One board per job. A single "best at AI" score hides the thing that matters: the card that wins at running a language model is often not the card that wins at upscaling, and the gap between them can be an order of magnitude.

What fits your card

Every model we have measured, grouped by the VRAM it actually needs. The figure is the lowest requirement any card recorded for that model, so it is the honest floor rather than a worst case.

Under 4GB runs on almost anything, including a GTX 1650 or a laptop chip

4 to 8GB RTX 3050, RTX 4060, RX 7600 class

8 to 12GB RTX 3060 12GB, RTX 4070, RX 7700 XT class

12 to 16GB RTX 4060 Ti 16GB, RTX 5070 Ti, RX 7800 XT class

16 to 24GB RTX 3090, RTX 4090, RTX 5090 class

Over 24GB datacenter only: A100, H100, H200, B200

Models by family

Every language model we have measured, grouped by the family it belongs to. Useful when you have picked a family and want to know which size to run.

Qwen 44 measured

Llama 11 measured

DeepSeek 10 measured

Gemma 10 measured

Dolphin 8 measured

Mistral 7 measured

Phi 6 measured

SmolLM 3 measured

Codestral 2 measured

GLM 2 measured

Hermes 2 measured

Devstral 1 measured

QwQ 1 measured

StarCoder 1 measured

Olmo 1 measured

Nemotron 1 measured

Other 11 measured

Boards by task

Text Generation

Running an LLM locally: tokens per second on the models people actually run.

1392 measured results across 121 models · fastest so far: NVIDIA GeForce RTX 5090

Image Generation

Turning a prompt into a picture, timed end to end at 1024px.

135 measured results across 11 models · fastest so far: NVIDIA B200

Video Generation

Rendering a clip from a prompt, measured in frames per second of output.

49 measured results across 2 models · fastest so far: NVIDIA B300

Fine-Tuning

Training throughput in tokens per second: LoRA on a fixed batch, the number that decides whether a fine-tune is an evening or a week.

43 measured results across 4 models · fastest so far: NVIDIA B300

LLM Serving

Throughput with 32 requests in flight at once, which is the question when a model sits behind an API rather than on your desk.

39 measured results across 4 models · fastest so far: NVIDIA B200

Image Editing

Changing an existing image rather than making a new one.

26 measured results across 2 models · fastest so far: NVIDIA B300

Depth Estimation

Reading 3D structure out of a flat photo, a stage in most 3D and relighting pipelines.

22 measured results across 2 models · fastest so far: NVIDIA RTX PRO 6000 Blackwell Workstation Edition

Segmentation

Isolating objects in an image, the groundwork for masking and compositing.

22 measured results across 2 models · fastest so far: NVIDIA B300

Vision Language

Describing and reasoning about a picture: captioning, OCR and visual question answering.

12 measured results across 2 models · fastest so far: NVIDIA H100 80GB HBM3

Background Removal

Cutting the subject out of a photo, the most common production task there is.

11 measured results across 1 model · fastest so far: NVIDIA H200

Upscaling

Enlarging an image without it falling apart.

11 measured results across 1 model · fastest so far: NVIDIA B300

Image to 3D

Turning a single picture into a 3D asset.

9 measured results across 2 models · fastest so far: NVIDIA H100 80GB HBM3

Image to Video

Animating a still you already have, the stage that turns keyframes into footage.

9 measured results across 3 models · fastest so far: NVIDIA H100 80GB HBM3

Text to Speech

Generating spoken audio, measured against realtime: 10x means ten seconds of speech every second.

8 measured results across 1 model · fastest so far: NVIDIA L40S

Music Generation

Generating music from a prompt, measured against realtime.

8 measured results across 1 model · fastest so far: NVIDIA H100 80GB HBM3

Speech to Text

Transcribing audio, measured against realtime, which is what decides whether a long recording is a coffee break or an afternoon.

8 measured results across 1 model · fastest so far: NVIDIA L40S

Speculative Decoding

Pairing a small draft model with a large one is sold as a free speedup. Measured per card, it is often a slowdown. Above 1.0x it helped, below 1.0x it cost you.

6 measured results across 1 model · fastest so far: NVIDIA L4