Skip to content
Model Pareto

Compare models

Side by side: quality per category, price, speed, and every benchmark cell with where it came from. Pick up to four.
Claude Opus 5.5Nex-N2.5-MiniSolar Mini 4

Overview

Lab
Anthropic
Nex AGI
Upstage
Released
Sep 22, 2026
Sep 8, 2026
Sep 22, 2026
Weights
Proprietary
Open
Proprietary
Input price
$/M tok
$4.00
$0.025Best
$0.10
Output price
$/M tok
$20
$0.10Best
$0.40
Blended price
$/M tok · 3:1 in:out
$8.00
$0.044Best
$0.175
Output speed
tok/s
92Best
—
82
Context
tokens
1MBest
262K
524K

Quality by category

0–100 within each category (100 = best tracked model). Rank is among all models ranked there. Click a row to focus it.

95.7±3.6Best
#1/119Mixedbest value
34.3±8.6
#67/119Lab-reportedbest value
Not ranked
88.3±4.3
#5/156Verified
Not ranked
Not ranked
94.0±4.8
#2/91Verifiedbest value
Not ranked
Not ranked
99.9±2.6Best
#1/200Verifiedbest value
Not ranked
55.4±5.5
#68/200Verified
99.6±4.1Best
#1/178Verifiedbest value
49.1±8.9
#38/178Lab-reportedbest value
Not ranked

Coding benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

SciCode
third-party
~#37/106~49.7%Estimated
estimated from related benchmarks
—

Writing & chat benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

Vision benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

Hard reasoning benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

Agents benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

~#2/5~58.6%Estimated
estimated from related benchmarks
~#4/5~48.5%Estimated
estimated from related benchmarks
—
Not comparable across labs (2)