Skip to content
Model Pareto

Compare models

Side by side: quality per category, price, speed, and every benchmark cell with where it came from. Pick up to four.
Claude Opus 5.5FLUX 3 VideoGranite 4.2 3B

Overview

Lab
Anthropic
Black Forest Labs
IBM Granite
Released
Sep 22, 2026
Jul 23, 2026
Aug 25, 2026
Weights
Proprietary
Proprietary
Open
Input price
$/M tok
$4.00
—
$0.03Best
Output price
$/M tok
$20
—
$0.12Best
Blended price
$/M tok · 3:1 in:out
$8.00
—
$0.052Best
Output speed
tok/s
92
—
220Best
Context
tokens
1MBest
—
131K
Price per second
1080p
—
$0.29
—

Quality by category

0–100 within each category (100 = best tracked model). Rank is among all models ranked there. Click a row to focus it.

96.9±3.4Best
#1/103Mixedbest value
Not ranked
17.2±8.0
#95/103Verifiedbest value
93.0±4.3
#3/146Verifiedbest value
Not ranked
Not ranked
94.5±4.8
#2/82Verifiedbest value
Not ranked
Not ranked
99.9±2.6Best
#1/182Verifiedbest value
Not ranked
23.4±3.7
#128/182Verified
99.6±4.1Best
#1/173Verifiedbest value
Not ranked
17.0±5.8
#140/173Verified
Not ranked
83.2±6.4
#8/34Verified
Not ranked

Coding benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

—
~last/15below measured range (<41.1)Estimated
estimated from related benchmarks

Writing & chat benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

Vision benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

Hard reasoning benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

Agents benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

~#2/12~84.6%Estimated
estimated from related benchmarks
—
~#2/5~58.7%Estimated
estimated from related benchmarks
—
~last/5below measured range (<47.5%)Estimated
estimated from related benchmarks
Not comparable across labs (1)

Video generation benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.