Upcoming models
This week4 launches due or landed
27 launches tracked
How this page works
Order
Rows are sorted by excitement (0–1) = impact² × timing × trust — impact counts double, so the most exciting models lead. Impact comes from the best expected rank (#1 → 1, #2–3 → 0.85, top 10 → 0.6, else 0.35), plus 0.15 when its expected quality tops today's #1 somewhere (▲), plus 0.04 for each extra category where it's likely top 3; halved for a re-release of a model we already rank. Without a rank it's 0.55 for a frontier lab (9 labs) and 0.3 otherwise. Timing is 0.15 + 0.55 × the chance it launches in the next 7 days + 0.30 × the chance in the next 30, taking the launch as equally likely on any day left in its window (a window that has already closed counts as the next 30 days plus twice the days it has slipped: long slips usually slip further). A model that's out but not yet ranked gets 1. Trust is 1 / 0.8 / 0.6 for official / rumour / leak, times 1 / 0.9 / 0.75 for high / medium / low confidence (0.8 / 0.5 for a medium / low-confidence leak), and × 0.8 when the newest source is more than 60 days old. Sources are listed best first: the lab itself, then top press, specialists, and aggregators last.
Sources
- OfficialThe developer announced the model or its timing.
- RumourReputable press or named insiders, not the developer.
- LeakUnverified signals: model ids in code or API lists, anonymous arena models, leaker accounts. Listed, never trusted — never high confidence.
Scanning
Every entry records when someone last checked its sources and searched for launch news. How soon it needs another look depends on where today falls relative to its expected window. The status is worked out in your browser with your local date, so it stays current between site updates.
- ReleasedIt has shipped but isn't on the leaderboard yetAdd now — it goes through intake the same day
- Window openToday is inside the expected launch windowEvery day
- ImminentThe window opens within 7 daysEvery 3 days
- LaterThe window is further outEvery 14 days for rumours and leaks, 30 for official dates
- Date passedThe window closed without a launchOnce right away, then every 7 days while it slips
- LaunchedThe model is on the leaderboardNever — it's tracked like any other model
Official means the developer announced the model or its timing; rumour means press reports, leaks or insiders only. Every source carries the sentence it was taken from. Maintainers get today's list with npm run upcoming.
Expected ranks
An expected rank is a prediction, not a score — a rough heuristic from history. We pair every model in our data with an earlier one from the same lab and modality, released one to eighteen months before it: the previous version in its series, or the lab's nearest earlier model in the same tier (price within 2×, or for open-weights models without a price, parameter count within 2×). The change in quality between the two, in each category, is one data point.
- Codingmedian +12.6 (+0.9 to +28.3) · 20 pairs
- Writing & chatmedian +11.7 (−9.9 to +29.7) · 39 pairs
- Visionmedian +10.0 (−7.8 to +27.3) · 23 pairs
- Hard reasoningmedian +11.7 (−10.3 to +30.8) · 52 pairs
- Agentsmedian +5.8 (−5.1 to +36.3) · 48 pairs
- Image generationmedian +14.5 (−6.5 to +33.9) · 11 pairs
- Video generationmedian +16.8 (−2.6 to +56.9) · 11 pairs
A category with 8 or more data points uses its own; otherwise it uses its modality's (text, image or video) and never borrows from another. Adding the median change to the previous model's quality and slotting the result into today's leaderboard gives the likely rank; the p10–p90 changes give the best and worst. In back-tests on our own data (leaving each pair out and predicting it from the rest) the range contained the actual rank 75% of the time (154 of 204).
Claims in the sources then narrow it: “beats X” keeps it at or above X, “matches X” puts the likely rank at X's, “trails X” keeps it below. Benchmark numbers from claims are shown but not used. Without a previous model, or with too little history, only claims place it. Each claim is marked by how far it can be trusted (official, rumour or leak), and a range shaped by a leaked claim is widened by one rank each way. Generations vary a lot, so treat the range as a hint about where to look, not a forecast.