Skip to content
Model Pareto

Compare models

Side by side: quality per category, price, speed, and every benchmark cell with where it came from. Pick up to four.
Gemini Omni 1.1 FlashGemini Omni FlashMolmo2-8BWan 3.0Up to 4 models — remove one to add another.

Overview

Lab
Google DeepMind
Google DeepMind
Ai2 (Allen Institute)
Alibaba Qwen
Released
Aug 27, 2026
May 19, 2026
Dec 11, 2025
Aug 19, 2026
Weights
Proprietary
Proprietary
Open
Proprietary
Context
tokens
—
—
37K
—
Price per second
1080p
$0.10Best
$0.10Best
—
$0.20

Quality by category

0–100 within each category (100 = best tracked model). Rank is among all models ranked there. Click a row to focus it.

Not ranked
Not ranked
18.5±5.1
#129/146Verified
Not ranked
Not ranked
Not ranked
17.2±8.2
#76/82Verified
Not ranked
Not ranked
Not ranked
13.4±5.8
#156/182Verified
Not ranked
Not ranked
Not ranked
13.9±10.0
#167/173Verified
Not ranked
94.3±6.3
#2/34Verified
96.9±3.0Best
#1/34Verifiedbest value
Not ranked
93.8±4.1
#3/34Verified

Writing & chat benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

Vision benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

MMMU
day 0
—
—
~#2/3~71.3%Estimated
estimated from related benchmarks
—
—
—
~last/9below measured range (<74.5%)Estimated
estimated from related benchmarks
—
—
—
~last/8below measured range (<75.8%)Estimated
estimated from related benchmarks
—
GDP.pdf (AA)
third-party
—
—
—
Chartography
third-party
—
—
~last/23below measured range (<9.0%)Estimated
estimated from related benchmarks
—

Hard reasoning benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

Agents benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

—
—
~#11/12~61.8%Estimated
estimated from related benchmarks
—
—
—
—
GDPval-AA (Elo)
third-party
—
—
—
—
—
~#80/80~0.2%Estimated
estimated from related benchmarks
—
—
—
~last/89below measured range (<3.3%)Estimated
estimated from related benchmarks
—

Video generation benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

VBench
day 0
—
—
—
—