Skip to content
Model Pareto

Compare models

Side by side: quality per category, price, speed, and every benchmark cell with where it came from. Pick up to four.
BorealClaude Fable 5.1Claude Opus 5.5Grok 4.7Up to 4 models — remove one to add another.

Overview

Lab
Creatify
Anthropic
Anthropic
SpaceXAI (xAI)
Released
Sep 15, 2026
Sep 1, 2026
Sep 22, 2026
Sep 21, 2026
Weights
Proprietary
Proprietary
Proprietary
Proprietary
Input price
$/M tok
—
$10
$4.00
$2.00Best
Output price
$/M tok
—
$50
$20
$6.00Best
Blended price
$/M tok · 3:1 in:out
—
$20
$8.00
$3.00Best
Output speed
tok/s
—
48
92Best
76
Context
tokens
—
1MBest
1MBest
500K
Price per second
1080p
$0.010
—
—
—

Quality by category

0–100 within each category (100 = best tracked model). Rank is among all models ranked there. Click a row to focus it.

Not ranked
86.2±3.6
#4/119Mixed
95.7±3.6Best
#1/119Mixedbest value
70.5±4.5
#15/119Verified
Not ranked
94.4±4.6Best
#2/156Verified
88.3±4.3
#5/156Verified
68.7±4.3
#30/156Verified
Not ranked
86.7±7.3
#4/91Verified
94.0±4.8Best
#2/91Verifiedbest value
64.6±4.8
#29/91Verified
Not ranked
95.7±3.7
#2/200Verified
99.9±2.6Best
#1/200Verifiedbest value
81.4±5.1
#16/200Verified
Not ranked
89.8±5.8
#2/178Verified
99.6±4.1Best
#1/178Verifiedbest value
87.1±7.8
#3/178Verifiedbest value
14.7±7.5
#34/36Verifiedbest value
Not ranked
Not ranked
Not ranked

Coding benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

—
~#2/14~93.5%Estimated
estimated from related benchmarks
~#2/14~94.7%Estimated
estimated from related benchmarks
~#2/14~85.9%Estimated
estimated from related benchmarks

Writing & chat benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

IFBench
day 0
—
~#2/134~82.8%Estimated
estimated from related benchmarks
~#2/134~82.8%Estimated
estimated from related benchmarks
~#26/134~72.9%Estimated
estimated from related benchmarks
MMMLU
day 0
—
~#2/5~92.4%Estimated
estimated from related benchmarks
~#2/5~92.3%Estimated
estimated from related benchmarks
~#2/5~89.6%Estimated
estimated from related benchmarks

Vision benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

MMMU
day 0
—
~#2/3~71.3%Estimated
estimated from related benchmarks
~#2/3~71.3%Estimated
estimated from related benchmarks
~#2/3~71.3%Estimated
estimated from related benchmarks
—
~#2/9~86.0%Estimated
estimated from related benchmarks
~#2/9~86.1%Estimated
estimated from related benchmarks
—
~#2/8~91.2%Estimated
estimated from related benchmarks
~#2/8~91.4%Estimated
estimated from related benchmarks
~#4/8~87.1%Estimated
estimated from related benchmarks

Hard reasoning benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

—
~#2/32~99.4%Estimated
estimated from related benchmarks
~#2/32~99.2%Estimated
estimated from related benchmarks

Agents benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

—
~#2/12~84.6%Estimated
estimated from related benchmarks
~#2/12~84.6%Estimated
estimated from related benchmarks
~#3/12~82.6%Estimated
estimated from related benchmarks
—
~#2/5~58.6%Estimated
estimated from related benchmarks
~#3/5~58.3%Estimated
estimated from related benchmarks
—
~#2/18~92.2%Estimated
estimated from related benchmarks
~#2/18~92.3%Estimated
estimated from related benchmarks
Not comparable across labs (1)

Video generation benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

—
—
—
~last/25below measured range (<1000)Estimated
estimated from related benchmarks
—
—
—
—
—
—
VBench
day 0
—
—
—
—