Skip to content
Model Pareto

Compare models

Side by side: quality per category, price, speed, and every benchmark cell with where it came from. Pick up to four.
Claude Opus 5.5GLM-5.3-MXFP4Nex-N2.5-Mini

Overview

Lab
Anthropic
Red Hat AI
Nex AGI
Released
Sep 22, 2026
Sep 15, 2026
Sep 8, 2026
Weights
Proprietary
Open
Open
Input price
$/M tok
$4.00
—
$0.025Best
Output price
$/M tok
$20
—
$0.10Best
Blended price
$/M tok · 3:1 in:out
$8.00
—
$0.044Best
Output speed
tok/s
92
—
—
Context
tokens
1MBest
1MBest
262K

Quality by category

0–100 within each category (100 = best tracked model). Rank is among all models ranked there. Click a row to focus it.

95.7±3.6Best
#1/119Mixedbest value
Not ranked
34.3±8.6
#67/119Lab-reportedbest value
88.3±4.3
#5/156Verified
Not ranked
Not ranked
94.0±4.8
#2/91Verifiedbest value
Not ranked
Not ranked
99.9±2.6Best
#1/200Verifiedbest value
66.1±8.3
#45/200Lab-reported
Not ranked
99.6±4.1Best
#1/178Verifiedbest value
Not ranked
49.1±8.9
#38/178Lab-reportedbest value

Coding benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

SciCode
third-party
—
~#37/106~49.7%Estimated
estimated from related benchmarks

Writing & chat benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

Vision benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

Hard reasoning benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

Agents benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

~#2/5~58.6%Estimated
estimated from related benchmarks
—
~#4/5~48.5%Estimated
estimated from related benchmarks
Not comparable across labs (2)