Skip to content
Model Pareto

Claude Opus 5.5

AnthropicProprietaryanthropic.com ↗
Input
$4.00/M tok
Output
$20/M tok
Blended
$8.00/M tok · 3:1
Speed
92tok/s
Context
1Mtokens

Speed measured on the xhigh effort variant.

Rankings

Benchmark scores

~ italic, dashed = no published score yet, estimated from related benchmarks.

Coding

Writing & chat

Vision

MMMUEstimated
~#2 of 3 (estimated)~71.3%
  1. #1Claude Haiku 4.573.2%
  2. ~#2Claude Opus 5.5Estimated71.3%
  3. #2Llama 4 Maverick69.4%
No published score — estimated from related benchmarks
~#2 of 9 (estimated)~86.1%
  1. #1Gemini 3.8 Flash86.2%
  2. ~#2Claude Opus 5.5Estimated86.1%
  3. #2Kimi K384.8%
No published score — estimated from related benchmarks
~#2 of 8 (estimated)~91.4%
  1. #1MiniMax-M391.6%
  2. ~#2Claude Opus 5.5Estimated91.4%
  3. #2Kimi K391.1%
No published score — estimated from related benchmarks

Hard reasoning

GPQA DiamondEstimated
~#2 of 173 (estimated)~96.1%
  1. #1GPT-6 Astra96.1%
  2. ~#2Claude Opus 5.5Estimated96.1%
  3. #2Gemini 3.8 Flash95.3%
No published score — estimated from related benchmarks
AIME (latest)Estimated
~#2 of 25 (estimated)~96.7%
  1. #1Inkling97.1%
  2. ~#2Claude Opus 5.5Estimated96.7%
  3. #2Muse Glimmer94.7%
No published score — estimated from related benchmarks

Agents

~#2 of 116 (estimated)~98.3%
  1. #1Step 3.7 Flash98.5%
  2. ~#2Claude Opus 5.5Estimated98.3%
  3. #2Gemini 3.1 Pro95.6%
No published score — estimated from related benchmarks
~#2 of 12 (estimated)~84.6%
  1. #1Kimi K384.8%
  2. ~#2Claude Opus 5.5Estimated84.6%
  3. #2Qwen3.8 27B84.3%
No published score — estimated from related benchmarks
OSWorld 2.0Estimated
~#2 of 5 (estimated)~58.7%
  1. #1GLM-5.3 Flash59.1%
  2. ~#2Claude Opus 5.5Estimated58.7%
  3. #2Kimi K358.3%
No published score — estimated from related benchmarks
BrowseCompEstimated
~#2 of 18 (estimated)~92.3%
  1. #1Atria Dawn Preview92.5%
  2. ~#2Claude Opus 5.5Estimated92.3%
  3. #2GPT-6 Astra91.5%
No published score — estimated from related benchmarks
~#3 of 89 (estimated)~50.0%
  1. #1Muse Spark 1.350.5%
  2. #2GLM-5.350.3%
  3. ~#3Claude Opus 5.5Estimated50.0%
  4. #3Qwen3.8 2.4T A95B49.1%
No published score — estimated from related benchmarks
Not comparable across labs (1)

Kept for reference, not counted in rankings: these use a lab-specific task set, answer key or scoring, so scores can't be compared fairly across labs. Only published scores are shown.