Qwen-Image-Edit · 47 cards measured first-party · Updated October 2026
Qwen-Image-Edit is the heaviest image workload in our suite: ~42GB of VRAM minimum, ~58GB to run clean, and it wants 58GB+ of system RAM on top. It excludes more hardware than anything except Llama 3.3 70B, and for the same reason. You can't optimise your way past a memory ceiling.
Benchmarked weights: Qwen/Qwen-Image-Edit

8.14 images/min on Qwen-Image-Edit, the ceiling. Measured on our bench. 288GB of VRAM, $40,000 at launch.
Best for: Qwen-Image-Edit work where you want the ceiling gone rather than the cheapest entry.

0.8 images/min on Qwen-Image-Edit, lowest launch price that still fits. Measured on our bench. 48GB of VRAM, $4,500 at launch.

2.64 images/min on Qwen-Image-Edit, most speed per dollar. Measured on our bench. 96GB of VRAM, $8,565 at launch. That is 0.31 images/min per $1,000 of launch price.
Compute-bound, but capacity is what decides whether you're in the conversation at all. The ~42GB floor puts this workload out of reach of every consumer card and most workstation cards, which is why the leaderboard below is almost entirely datacenter silicon.
Qwen-Image-Edit: speed on every GPU we have data for
Top 15 shown; 4 more cards in the full table below.
Single stream, batch size 1. 36 of the 58 cards on this page were measured first-party by us; the rest are anchored estimates against those measurements and are labelled in the table below.
Efficiency: images/min per 100W drawn
Power is the average pulled during the run, sampled at 1Hz. The fastest card is often not the one here, and for anything left running this is the number that shows up on the bill.
Value: images/min per $1,000 of MSRP
Launch price, not street price, so it ages. A speed leaderboard always crowns the most expensive card; this is the counterweight.
Won't fit, Qwen-Image-Edit gates these cards outright
| GPU | VRAM | Why it fails |
|---|---|---|
| NVIDIA A100 40GB PCIe | 40GB | Needs ~42GB VRAM |
| NVIDIA A100 40GB SXM4 | 40GB | Needs ~42GB VRAM |
| NVIDIA GeForce RTX 5090 | 32GB | Needs ~42GB VRAM |
| NVIDIA RTX 5000 Ada Generation | 32GB | Needs ~42GB VRAM |
| NVIDIA RTX PRO 4500 Blackwell | 32GB | Needs ~42GB VRAM |
| NVIDIA GeForce RTX 3090 Ti | 24GB | Needs ~42GB VRAM |
| NVIDIA GeForce RTX 3090 | 24GB | Needs ~42GB VRAM |
| NVIDIA GeForce RTX 4090 | 24GB | Needs ~42GB VRAM |
| NVIDIA A10G | 24GB | Needs ~42GB VRAM |
| NVIDIA L4 | 24GB | Needs ~42GB VRAM |
| NVIDIA Quadro RTX 6000 (Turing) | 24GB | Needs ~42GB VRAM |
| NVIDIA RTX 4500 Ada Generation | 24GB | Needs ~42GB VRAM |
| NVIDIA RTX A5000 | 24GB | Needs ~42GB VRAM |
| NVIDIA RTX A5500 | 24GB | Needs ~42GB VRAM |
| NVIDIA RTX PRO 4000 Blackwell | 24GB | Needs ~42GB VRAM |
| NVIDIA Titan RTX | 24GB | Needs ~42GB VRAM |
| NVIDIA RTX 4000 (Ada Generation) | 20GB | Needs ~42GB VRAM |
| NVIDIA RTX A4500 | 20GB | Needs ~42GB VRAM |
| NVIDIA GeForce RTX 4060 Ti 16GB | 16GB | Needs ~42GB VRAM |
| NVIDIA GeForce RTX 4070 Ti Super | 16GB | Needs ~42GB VRAM |
| GeForce RTX 4080 Super | 16GB | Needs ~42GB VRAM |
| NVIDIA GeForce RTX 4080 | 16GB | Needs ~42GB VRAM |
| GeForce RTX 5060 Ti | 16GB | Needs ~42GB VRAM |
| GeForce RTX 5070 Ti | 16GB | Needs ~42GB VRAM |
| GeForce RTX 5080 | 16GB | Needs ~42GB VRAM |
| NVIDIA Quadro RTX 5000 | 16GB | Needs ~42GB VRAM |
| NVIDIA RTX 2000 Ada Generation | 16GB | Needs ~42GB VRAM |
| NVIDIA RTX A4000 | 16GB | Needs ~42GB VRAM |
| NVIDIA T4 | 16GB | Needs ~42GB VRAM |
| NVIDIA GeForce RTX 3060 | 12GB | Needs ~42GB VRAM |
| NVIDIA GeForce RTX 3080 Ti | 12GB | Needs ~42GB VRAM |
| NVIDIA GeForce RTX 4070 Super | 12GB | Needs ~42GB VRAM |
| NVIDIA GeForce RTX 4070 Ti | 12GB | Needs ~42GB VRAM |
| NVIDIA GeForce RTX 4070 | 12GB | Needs ~42GB VRAM |
| GeForce RTX 5070 | 12GB | Needs ~42GB VRAM |
| NVIDIA TITAN V | 12GB | Needs ~42GB VRAM |
| NVIDIA TITAN X (Pascal) | 12GB | Needs ~42GB VRAM |
| NVIDIA TITAN Xp | 12GB | Needs ~42GB VRAM |
| GeForce GTX 1080 Ti | 11GB | Needs ~42GB VRAM |
| NVIDIA GeForce RTX 2080 Ti Founders Edition | 11GB | Needs ~42GB VRAM |
| NVIDIA GeForce RTX 3080 | 10GB | Needs ~42GB VRAM |
| NVIDIA GeForce GTX 1070 Ti | 8GB | Needs ~42GB VRAM |
| NVIDIA GeForce GTX 1080 | 8GB | Needs ~42GB VRAM |
| NVIDIA GeForce RTX 2060 Super | 8GB | Needs ~42GB VRAM |
| NVIDIA GeForce RTX 2070 SUPER | 8GB | Needs ~42GB VRAM |
| NVIDIA GeForce RTX 2070 | 8GB | Needs ~42GB VRAM |
| NVIDIA GeForce RTX 2080 Super | 8GB | Needs ~42GB VRAM |
| NVIDIA GeForce RTX 2080 Founders Edition | 8GB | Needs ~42GB VRAM |
| NVIDIA GeForce RTX 3050 | 8GB | Needs ~42GB VRAM |
| NVIDIA GeForce RTX 3060 Ti | 8GB | Needs ~42GB VRAM |
| NVIDIA GeForce RTX 3070 Ti | 8GB | Needs ~42GB VRAM |
| NVIDIA GeForce RTX 3070 Founders Edition | 8GB | Needs ~42GB VRAM |
| GeForce RTX 4060 | 8GB | Needs ~42GB VRAM |
| NVIDIA GeForce RTX 5050 | 8GB | Needs ~42GB VRAM |
| NVIDIA GeForce RTX 5060 | 8GB | Needs ~42GB VRAM |
| NVIDIA GeForce GTX 1660 Super | 6GB | Needs ~42GB VRAM |
| NVIDIA GeForce GTX 1660 Ti | 6GB | Needs ~42GB VRAM |
| NVIDIA GeForce GTX 1660 | 6GB | Needs ~42GB VRAM |
| NVIDIA GeForce RTX 2060 | 6GB | Needs ~42GB VRAM |
| NVIDIA RTX A2000 | 6GB | Needs ~42GB VRAM |
Showing 14 of 38. No driver update fixes a VRAM ceiling.
Full Qwen-Image-Edit leaderboard, every card that runs it
| GPU | Result | VRAM | Source |
|---|---|---|---|
| NVIDIA B300 | 8.14 images/min | 288GB | Measured |
| NVIDIA B200 | 4.86 images/min | 192GB | Measured |
| NVIDIA B100 | 3.88 images/min | 192GB | Estimated |
| NVIDIA GH200 Grace Hopper | 3.76 images/min | 141GB | Estimated |
| NVIDIA H200 | 3.76 images/min | 141GB | Measured |
| NVIDIA H100 NVL | 3.68 images/min | 94GB | Estimated |
| NVIDIA H100 80GB HBM3 | 3.68 images/min | 80GB | Measured |
| NVIDIA H800 80GB | 3.68 images/min | 80GB | Estimated |
| NVIDIA H100 PCIe | 3.18 images/min | 80GB | Estimated |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 2.64 images/min | 96GB | Measured |
| NVIDIA RTX PRO 6000 Blackwell Server Edition | 2.62 images/min | 96GB | Measured |
| NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition | 2.06 images/min | 96GB | Estimated |
| NVIDIA A100 80GB SXM4 | 1.64 images/min | 80GB | Measured |
| NVIDIA A800 80GB | 1.64 images/min | 80GB | Estimated |
| NVIDIA A100 80GB PCIe | 1.54 images/min | 80GB | Measured |
| NVIDIA L40S | 1.06 images/min | 48GB | Measured |
| NVIDIA RTX PRO 5000 Blackwell | 0.8 images/min | 48GB | Measured |
| NVIDIA L40 | 0.48 images/min | 48GB | Measured |
| NVIDIA RTX 5880 Ada Generation | 0.17 images/min | 48GB | Estimated |
Tap any column to sort. Measured = we rented and ran this card ourselves. Estimated = interpolated against our measured anchors, never blended silently.
Because this workload is tensor-compute bound, the ranking tracks architecture generation and tensor throughput rather than memory bandwidth, the reverse of our LLM charts. The same two cards can swap places entirely depending on which of these pages you're reading. That's the reason we run twelve workloads instead of publishing one score. A GPU isn't fast or slow. It's fast at some things and gated out of others, and which of those matters depends entirely on what you're actually going to run.
NVIDIA B300 tops our Qwen-Image-Edit leaderboard at 8.14 images/min (measured), 4688% of the way clear of the slowest card that still fits. But the number that decides most purchases isn't on the chart. It's the 60 cards that can't run Qwen-Image-Edit at all. This is a compute workload: buy architecture generation, not raw VRAM, as long as you clear the floor first.
Every ranking on this page comes from our own benchmark runs, not vendor claims. Cards marked Measured were rented and run by us; cards marked Estimated are interpolated per workload against those measured anchors and are labelled on every row, we never blend the two silently. LLMs run on llama.cpp (llama-bench) at Q4_K_M with -p 512 -n 128. Diffusion and video run on diffusers/ComfyUI at BF16, with SDXL at FP16. Each workload gets a warmup pass plus multiple timed runs (5 for small LLMs, 3 for large models and images, 2 for video); we publish the mean as the result and the minimum as the 1% low. Run-to-run variance is under 0.5%. Telemetry, power, temperature, utilisation, clocks, peak VRAM, is sampled at 1 Hz for the duration of every run. Where a model exceeds a card's VRAM we publish a hard won't-fit result rather than quietly dropping to a smaller quantisation. A card that can't run a model scores zero on it. Silently swapping precision to make a number appear would make every number on this site meaningless. All figures are single-GPU, single-stream, batch-size-1. That is the honest way to measure what one card does for one user, and it is deliberately not how a datacenter serves a model. Vendor and MLPerf figures use large batches across many GPUs and will be far higher. Neither is wrong, they answer different questions. Ours answers 'what will this card do for me'.