How GPU Battle works · every number, sourced · Updated July 2026

How We Benchmark: Measured First-Party, Estimated Honestly

Every number on GPU Battle is one of two things: a first-party measurement from our own bench, or an anchored estimate that is labeled as one. This page explains exactly how each is produced, what the badges mean, and the automated checks every data point passes before it is published. If you ever find a number here that can't be traced to this methodology, that's a bug, and we treat it like one.

157
GPUs in the database
desktop, laptop, workstation, datacenter
51
GPUs measured first-party on our AI suite
rented, benched, logged
21
GPUs measured first-party on our gaming bench
one CPU, identical settings, captures kept
284
first-party gaming records
1080p and 1440p, every run screenshotted

The three tiers of data, and the badges that mark them.Measured means we ran the workload ourselves on that exact card and logged the result. For gaming, every measured run has a benchmark capture you can view on the card's page. For AI, every measured run carries its harness version, driver, run date, per-run timings and run-to-run variance. Est. means an anchored estimate: we interpolate the card against cards we *did* measure, using its position on well-established performance ladders. Estimates exist so you can still compare a card we haven't benched yet. But they are labeled everywhere they appear, they never count as measurements, and comparison pages built only on estimates are deliberately kept out of Google's index. A small number of cards sit in a third tier: cards we measured at 1080p whose 1440p figures are derived from our own 1080p runs, scaled against a card we measured at 1440p on the same bench. Those pages keep the raw 1080p measurements visible, so nothing is hidden behind the derivation.

Gaming: one suite, identical settings, or it doesn't count. Our gaming figures come from a fixed game suite run at 1080p and 1440p with ray tracing off, frame generation off, and upscaling off: all on one personal bench: an AMD Ryzen 7 7800X3D system, the same CPU for every card. For every card and every game we publish the average FPS, the 1% low, frametime, VRAM used, board power and peak temperature. We never mix presets or resolutions inside a ranking. That's the single most common way GPU comparisons quietly become meaningless. One bench, one ladder. Every first-party gaming number on this site comes from that one bench. It is the standing platform: as cards are added they run the same suite on the same machine and join the same ladder, so every row stays comparable to every other. The 1% low is published next to every average on purpose. The average is the number everyone quotes; the 1% low is the number you feel.

AI: 12 workloads, rented hardware, real logs. We rent the actual cards across three cloud providers, Vast.ai, RunPod and Modal, depending on which has the silicon (they all have their pros and cons), and run a pinned 12-workload suite on each: an LLM ladder from Qwen3 4B up to Llama 3.3 70B (llama.cpp, Q4_K_M), image generation (SDXL, Z-Image Turbo, FLUX.1-dev), image editing (FLUX Kontext, Qwen-Image-Edit), and video generation (LTX-Video, Wan 2.2). The harness samples power and temperature at 1Hz through every run and records tokens-per-watt, peak VRAM, and per-run variance. "Won't fit" is a result, not a gap. When a model exceeds a card's VRAM at the tested precision, we publish that as a hard gate instead of quietly dropping to a smaller quantization. A model that doesn't fit doesn't run slowly, it doesn't run. That's information you need before buying, so it counts against the card's score.

The AI Score, exactly as computed. AI Score is 100 × the geometric mean, across all 12 workloads, of each result normalized against the best result any card achieved on that workload. A workload that won't fit scores a floor value of 0.01, a hard penalty that makes the VRAM ceiling part of the ranking. A workload we haven't run yet on a card that *would* fit is imputed neutrally from the card's tested results, adjusted for how hard that workload is for every other card. It can't inflate a card's score above its evidence. Every card's page records how many of its 12 slots were tested, gated, or imputed. When we re-run a card or add measurements, scores are recomputed for the whole board at once, with one formula, never hand-adjusted.

VR figures are estimates today, labeled as such. Our VR numbers are formula-based estimates anchored to published FCAT-VR measurements, pending our own headset test rig. They are marked estimated on every page, and VR-only comparison pages are excluded from search indexing entirely until they're backed by measurements.

The validation gate: what runs before anything ships. Every data change passes an automated validation suite before deployment. It checks physical plausibility (a 1% low can't exceed an average; VRAM used can't exceed VRAM fitted), cross-checks specs against every vertical of the database, verifies VRAM gates are consistent across cards (a 24GB card can't be 'unable to fit' a model a 12GB card ran), and runs tier-monotonicity checks, within a product family, a faster card outscoring a slower one is asserted, not assumed. A deploy with a failing check is blocked. We built this because we found errors in our own data and decided the honest fix was structural, not cosmetic. The same suite that caught them now guards every update.

Our verdict

The short version: green check means we ran it and kept the receipts; Est. means we're telling you it's an interpolation; and every ranking on this site is computed from one formula over one suite, gated by automated checks. When a number changes, it's because the data changed, and the methodology on this page is the whole story of where every figure comes from.

FAQ

Why publish estimates at all?
Because refusing to compare a card we haven't benched yet helps nobody. An anchored estimate with a visible label is more honest than silence and far more honest than an unlabeled guess. Estimates are also structurally quarantined: they never count as measurements, and estimate-only comparison pages are excluded from search indexing.
Why did a score or FPS number change since I last looked?
Three legitimate reasons: we measured the card (estimate replaced by measurement), we measured more cards (estimates are re-anchored against a bigger measured set), or the whole board was recomputed after a methodology fix. Scores are never hand-edited, every change traces to a data change.
What does 'won't fit' mean on an AI page?
The model exceeds the card's VRAM at our tested precision, so it cannot load, we record the VRAM it would need and score the workload at the floor. We publish these instead of silently switching to a smaller quantization, because a smaller quant is a different benchmark.
How is this different from aggregated benchmark sites?
Aggregators average user submissions from thousands of uncontrolled machines: different CPUs, drivers, settings and background load. Every measured number here comes from one controlled setup per suite, and every estimated number says it's estimated. Fewer data points, but every one is comparable to every other.
What hardware does the bench run on?
Gaming: a personal AMD Ryzen 7 7800X3D test rig, identical settings across all cards (this is the standing gaming bench; every first-party gaming number on the site comes from it, and new cards are added to the same ladder). AI: rented cloud instances of the actual cards across Vast.ai, RunPod and Modal, with the harness pinned to specific driver, backend, and model versions per suite, each card's page lists the exact versions used for its run.
Can I see the raw evidence for a number?
Measured gaming runs keep their benchmark captures, viewable on each card's page. Measured AI runs publish per-run timings, variance, power, thermals and peak VRAM in the card's benchmark log. If a number has no evidence trail, it's labeled an estimate.
Why are only 21 cards measured for gaming when 51 are measured for AI?
Because AI cards are rented and gaming cards are owned. An H200 can be benched for a few dollars an hour on RunPod or Vast; a gaming card has to be physically bought, installed and run through the suite by hand. So the AI ladder grows with budget and the gaming ladder grows with hardware. The 21 gaming cards are the ones that have physically been in the bench, and that number goes up as cards are added, not by aggregating anyone else's numbers.