Skip to content
Model Pareto

Compare models

Side by side: quality per category, price, speed, and every benchmark cell with where it came from. Pick up to four.
Agnes-Video-2.5Claude Opus 5.5Granite 4.2 3B

Overview

Lab
Sapiens AI
Anthropic
IBM Granite
Released
Aug 22, 2026
Sep 22, 2026
Aug 25, 2026
Weights
Proprietary
Proprietary
Open
Input price
$/M tok
—
$4.00
$0.03Best
Output price
$/M tok
—
$20
$0.12Best
Blended price
$/M tok · 3:1 in:out
—
$8.00
$0.052Best
Output speed
tok/s
—
92
220Best
Context
tokens
—
1MBest
131K
Price per second
1080p
$0.025
—
—

Quality by category

0–100 within each category (100 = best tracked model). Rank is among all models ranked there. Click a row to focus it.

Not ranked
96.6±3.5Best
#1/105Mixedbest value
17.4±8.0
#97/105Verifiedbest value
Not ranked
93.0±4.3
#3/147Verifiedbest value
Not ranked
Not ranked
94.5±4.8
#2/83Verifiedbest value
Not ranked
Not ranked
99.9±2.6Best
#1/184Verifiedbest value
23.3±3.7
#129/184Verified
Not ranked
99.6±4.1Best
#1/175Verifiedbest value
17.1±5.8
#142/175Verified
44.2±9.0
#26/36Verifiedbest value
Not ranked
Not ranked

Coding benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

—
~last/15below measured range (<41.1)Estimated
estimated from related benchmarks

Writing & chat benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

Vision benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

Hard reasoning benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

Agents benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

—
~#2/12~84.6%Estimated
estimated from related benchmarks
—
~#2/5~58.7%Estimated
estimated from related benchmarks
~last/5below measured range (<47.5%)Estimated
estimated from related benchmarks
Not comparable across labs (1)

Video generation benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.