Fast enough to be useful, honest about what we know.
The 0–2 day promise
A new model appears on the leaderboard within 0–2 days of launch. On day zero we proxy the numbers the lab publishes in its model card or launch post. Over the following one to fourteen days, third-party evaluators — Artificial Analysis, LMArena and the benchmark maintainers — publish independent results, and each lab-reported score is upgraded to a verified one.
- 01 · Day 0Lab-reportedLaunch
Model-card numbers ingested. Gaps estimated. Ranked.
- 02 · Days 1–3Few votesFirst third-party runs
Artificial Analysis price, speed and indices land.
- 03 · Days 3–14VerifiedVerified
Arena votes settle; reproductions replace lab claims.
Where scores come from
Every score, and every model's composite, carries a badge. Colour is never the only signal — each tier has its own icon and label.
- Lab-reportedTaken from the model card or launch post. Not yet independently reproduced.
- VerifiedIndependently reproduced by a third party (Artificial Analysis, LMArena or the benchmark maintainers).
- Few votesEarly arena result with few votes so far; the score may still shift. Unlike Estimated, this is a real measured score; it just isn't settled yet.
- EstimatedNo published score. Estimated from this model's results on related benchmarks.
- MixedSome scores are verified, others still lab-reported. Used for a model's composite, never a single score.
Lab-reported numbers are not independently checked: labs choose their own scaffolds, sampling and effort settings. We record those settings where they are disclosed and prefer verified results the moment they exist.
Putting benchmarks on one scale
Benchmarks live on different scales — percentages, Elo ratings, composite indices — and differ in how saturated they are. Before averaging, each benchmark is normalized to 0–100 among the models we track (percentile or percent-of-best), so a point on a hard benchmark counts as much as a point on an easy one. A category's quality is the weighted mean of its normalized benchmarks.
Filling gaps (estimated scores)
When a model has no published score on a benchmark yet, we estimate one from how it did on related benchmarks — and always label it as an estimate.
On launch day only some benchmarks are published. Rather than ranking a model on two benchmarks and its rival on five, each missing score is predicted from the benchmarks that correlate with it across the models we track (a per-pair linear fit, pulled toward the model's own average when the fit is weak). Estimates count for less than measured scores, are never stored as data, are recomputed on every build and are always shown italic and dashed with an Estimated badge.
Models that are strong at coding tend to be strong at reasoning, agents, writing and vision too, so for the text categories an estimate also uses how the model scored in the other categories (measured scores only). A model with one vision benchmark but solid coding and reasoning results is estimated from all of that, not from one number. A model with no benchmark at all in a category, but measured in at least two others, can still be listed there, marked Estimated from other categories.
The range shown around each score is measured, not guessed: we take well-tested models, hide all but one, two or three of their scores, re-estimate them, and size the range so that about 8 in 10 of those estimates land inside it, separately for each category. Fully measured models get a tight range; mostly estimated ones a wider one.
Labs publish their best results, so are estimates for the benchmarks a lab left out too generous? We checked: for 32 models where an independent tester later measured a benchmark the lab had skipped, our estimate from the lab's own numbers was on average 0.7 points (of 100) below the measured score (95% range −3.2 to +1.8, over- and underestimates equally common). We found no optimism to correct, so estimates are not adjusted and the range is symmetric. Benchmark variants that only one lab runs are shown for reference but never used to estimate anything.
Each model shows how much of its category is actually measured. On the charts, a model with fewer than half its benchmarks measured is drawn as a pale, outlined dot; one estimated purely from other categories as a hollow, dashed dot. Open-weights models with no API price are plotted at their self-hosting estimate as a diamond. Models with neither a price nor an estimate, or no speed data, sit in a narrow strip at the left edge of the chart, at their quality. Neither is ever on the frontier.
Best-value (Pareto) frontier
A model is on the frontier when no other model is both better and cheaper (or better and faster). Those are the only models worth choosing on these axes; everything below the line is dominated. Price is plotted on a log scale because prices span three orders of magnitude.
Price & speed
Price and speed come from Artificial Analysis and are normalized per modality:
| Modality | Price | Speed |
|---|---|---|
| Text | $/M tokens, blended 3:1 input:output | Output tokens/s |
| Image | $/image | Seconds per image (lower is better) |
| Video | $/second at 1080p | Render seconds per clip (lower is better) |
Estimated speed. When a model has no measured speed yet, we estimate one so it can still sit on the speed chart. Open-weights models are estimated from their size (the parameters used per token: smaller models are generally faster); other models from a sibling model at a similar price, or else from their price and lab. Image models are estimated from their price per image; for video models nothing we know (price, lab, sibling) predicts render time better than a plain guess, so they get the typical render time of the video models that are measured, with a wide range (about 40–170 s). Each estimate is fitted on the models that do have a measured speed and is typically off by a factor of about 1.4–2, so we show a range that holds the real value for about 8 in 10 measured models. Estimated speeds are drawn as ◇, shown as “~80” with an “est.” tag, and never put a model on the frontier.
Self-hosting estimates
For open-weights models we estimate what it would cost to run them yourself on rented GPUs, so models that no API sells can still be placed on the price chart. It is an estimate, not a quote, and is always shown as a range in muted italic with a diamond marker.
- Hardware from parameter count. The model's total parameters decide how much GPU memory the weights need at a given precision (FP8, BF16 or int4), plus room for the working memory of several requests at once. We pick the smallest standard setup that fits, for example two H100 80GB cards.
- Throughput from active parameters. Mixture-of-experts models only use part of their weights per token, so they generate faster than their size suggests. Throughput is for many requests batched together, not a single chat.
- Price from GPU rental. The hourly price of that hardware at typical cloud rates, divided by the tokens it can produce in an hour, blended 3:1 input:output like API prices.
- Calibrated against real hosts. Where third-party providers sell the same open model, we compare and adjust the estimate so it lands in the same range as what they actually charge.
- The range is busy vs light use. The low end assumes the GPUs are kept busy with batched traffic; the high end assumes light, bursty use where you pay for idle time. The chart plots the middle.
Where a model is small enough to run on one consumer GPU or a Mac (usually at int4), we say so. Each model's page lists the exact assumptions behind its estimate. Self-host estimates never put a model on the best-value frontier, and you can hide them with the API prices switch above the chart.
Benchmark portfolio
Each category combines benchmarks that labs publish on launch day with third-party evaluations that arrive later. Day-0 benchmarks make the 0–2 day promise possible; third-party ones keep it honest.
| Category | Lab-reported Day 0 · model cards | Verified Days 1–14 · third-party |
|---|---|---|
| Coding | SWE-bench Verified, Terminal-Bench | LiveCodeBench, AA Coding Index |
| Writing & chat | IFEval, MMMLU | LMArena text + per-language, EQ-Bench |
| Vision | MMMU / MMMU-Pro | OCR / document benchmarks |
| Hard reasoning | GPQA Diamond, HLE, AIME | AA Intelligence Index |
| Agents | τ²-bench, OSWorld, BrowseComp | AA agentic evals |
| Image generation | GenEval, DPG-Bench (open models) | AA Image Arena, LMArena T2I & edit |
| Video generation | VBench (occasionally) | AA Video Arena, LMArena video |
Arena ratings with under 4,000 votes are marked Few votes: a real, measured result that may still shift as more votes come in.