How GPU Battle works · every number, sourced · Updated July 2026
Every number on GPU Battle is one of two things: a first-party measurement from our own bench, or an anchored estimate that is labeled as one. This page explains exactly how each is produced, what the badges mean, and the automated checks every data point passes before it is published. If you ever find a number here that can't be traced to this methodology, that's a bug, and we treat it like one.
The three tiers of data, and the badges that mark them. ✓ Measured means we ran the workload ourselves on that exact card and logged the result. For gaming, every measured run has a benchmark capture you can view on the card's page. For AI, every measured run carries its harness version, driver, run date, per-run timings and run-to-run variance. Est. means an anchored estimate: we interpolate the card against cards we *did* measure, using its position on well-established performance ladders. Estimates exist so you can still compare a card we haven't benched yet. But they are labeled everywhere they appear, they never count as measurements, and comparison pages built only on estimates are deliberately kept out of Google's index. A small number of cards sit in a third tier: cards we measured at 1080p whose 1440p figures are derived from our own 1080p runs, scaled against a card we measured at 1440p on the same bench. Those pages keep the raw 1080p measurements visible, so nothing is hidden behind the derivation.
Gaming: one suite, identical settings, or it doesn't count. Our gaming figures come from a fixed game suite run at 1080p and 1440p with ray tracing off, frame generation off, and upscaling off: all on one personal bench: an AMD Ryzen 7 7800X3D system, the same CPU for every card. For every card and every game we publish the average FPS, the 1% low, frametime, VRAM used, board power and peak temperature. We never mix presets or resolutions inside a ranking. That's the single most common way GPU comparisons quietly become meaningless. One bench, one ladder. Every first-party gaming number on this site comes from that one bench. It is the standing platform: as cards are added they run the same suite on the same machine and join the same ladder, so every row stays comparable to every other. The 1% low is published next to every average on purpose. The average is the number everyone quotes; the 1% low is the number you feel.
AI: 12 workloads, rented hardware, real logs. We rent the actual cards across three cloud providers, Vast.ai, RunPod and Modal, depending on which has the silicon (they all have their pros and cons), and run a pinned 12-workload suite on each: an LLM ladder from Qwen3 4B up to Llama 3.3 70B (llama.cpp, Q4_K_M), image generation (SDXL, Z-Image Turbo, FLUX.1-dev), image editing (FLUX Kontext, Qwen-Image-Edit), and video generation (LTX-Video, Wan 2.2). The harness samples power and temperature at 1Hz through every run and records tokens-per-watt, peak VRAM, and per-run variance. "Won't fit" is a result, not a gap. When a model exceeds a card's VRAM at the tested precision, we publish that as a hard gate instead of quietly dropping to a smaller quantization. A model that doesn't fit doesn't run slowly, it doesn't run. That's information you need before buying, so it counts against the card's score.
The AI Score, exactly as computed. AI Score is 100 × the geometric mean, across all 12 workloads, of each result normalized against the best result any card achieved on that workload. A workload that won't fit scores a floor value of 0.01, a hard penalty that makes the VRAM ceiling part of the ranking. A workload we haven't run yet on a card that *would* fit is imputed neutrally from the card's tested results, adjusted for how hard that workload is for every other card. It can't inflate a card's score above its evidence. Every card's page records how many of its 12 slots were tested, gated, or imputed. When we re-run a card or add measurements, scores are recomputed for the whole board at once, with one formula, never hand-adjusted.
VR figures are estimates today, labeled as such. Our VR numbers are formula-based estimates anchored to published FCAT-VR measurements, pending our own headset test rig. They are marked estimated on every page, and VR-only comparison pages are excluded from search indexing entirely until they're backed by measurements.
The validation gate: what runs before anything ships. Every data change passes an automated validation suite before deployment. It checks physical plausibility (a 1% low can't exceed an average; VRAM used can't exceed VRAM fitted), cross-checks specs against every vertical of the database, verifies VRAM gates are consistent across cards (a 24GB card can't be 'unable to fit' a model a 12GB card ran), and runs tier-monotonicity checks, within a product family, a faster card outscoring a slower one is asserted, not assumed. A deploy with a failing check is blocked. We built this because we found errors in our own data and decided the honest fix was structural, not cosmetic. The same suite that caught them now guards every update.
The short version: green check means we ran it and kept the receipts; Est. means we're telling you it's an interpolation; and every ranking on this site is computed from one formula over one suite, gated by automated checks. When a number changes, it's because the data changed, and the methodology on this page is the whole story of where every figure comes from.