Claude Opus 5
- Input
- $5.00/M tok
- Output
- $25/M tok
- Blended
- $10/M tok · 3:1
- Speed
- 54tok/s
- Context
- 1Mtokens
Previous generation (superseded by Opus 5.5).
Rankings
Coding
Mixed#5 / 103rank by quality83.3±3.2quality / 1006/7 benchmarks measuredWriting & chat
Verified#2 / 146rank by quality93.4±4.3quality / 1003/5 benchmarks measuredPrice frontierVision
Verified#9 / 82rank by quality77.4±4.8quality / 1003/6 benchmarks measuredHard reasoning
Verified#5 / 182rank by quality91.7±3.7quality / 1003/4 benchmarks measuredAgents
Mixed#8 / 173rank by quality84.8±4.6quality / 1004/7 benchmarks measured
Benchmark scores
~ italic, dashed = no published score yet, estimated from related benchmarks.
Coding
SWE-bench Pro (Public)Lab-reported
#4 of 1879.2%
- #1Claude Opus 5.589.9%
- #2Claude Sonnet 5.581.3%
- #3Claude Fable 5.181.2%
- #4Claude Opus 579.2%
- #5Intern-S2-397B68.5%
anthropic.com · 2026-07-24
Terminal-Bench 4.0Verified
#5 of 8349.0%
- #1Claude Sonnet 5.563.6%
- #2Claude Opus 5.559.6%
- #4Claude Fable 5.152.0%
- #5Claude Opus 549.0%
- #6GPT-6 Sol43.9%
artificialanalysis.ai · 2026-09-24
LiveCodeBenchEstimated
~#3 of 20 (estimated)~91.3%
No published score — estimated from related benchmarks
SciCodeVerified
#14 of 9556.4%
- #1Claude Opus 5.566.9%
- #2Claude Fable 5.163.1%
- #13GPT-6 Astra56.5%
- #14Claude Opus 556.4%
- #15GPT-5.555.8%
artificialanalysis.ai · 2026-09-24
Writing & chat
LMArena Text, Non-English (Elo)Verified
#4 of 531483 Elo
- #1Claude Opus 5.51501
- #2Claude Fable 5.1Few votes1496
- #3Muse Spark 1.3Few votes1487
- #4Claude Opus 51483
- #4Gemini 3.8 FlashFew votes1483
lmarena.ai · 2026-09-25
#3 of 252133 Elo
eqbench.com · 2026-09-24
Vision
CharXiv ReasoningEstimated
~#2 of 9 (estimated)~84.9%
No published score — estimated from related benchmarks
OmniDocBench v1.5Estimated
~#4 of 8 (estimated)~90.7%
- #1MiniMax-M391.6%
- #2Kimi K391.1%
- #2Qwen3.8 27B91.1%
- ~#4Claude Opus 5Estimated90.7%
- #4Gemini 3.1 Pro85.3%
No published score — estimated from related benchmarks
Hard reasoning
GPQA DiamondVerified
#11 of 17293.2%
- #1GPT-6 Astra96.1%
- #2Gemini 3.8 Flash95.3%
- #5Step 5 Preview93.5%
- #11Claude Opus 593.2%
- #12MiniMax-M392.9%
artificialanalysis.ai · 2026-09-24
Humanity's Last ExamVerified
#4 of 17854.9%
- #1Claude Opus 5.561.4%
- #2Claude Fable 5.159.1%
- #3Claude Sonnet 5.555.0%
- #4Claude Opus 554.9%
- #5GPT-6 Astra54.7%
artificialanalysis.ai · 2026-09-24
AIME (latest)Estimated
~#2 of 25 (estimated)~96.9%
No published score — estimated from related benchmarks
Agents
τ²-bench (Telecom)Estimated
~#2 of 116 (estimated)~98.3%
No published score — estimated from related benchmarks
OSWorld-VerifiedEstimated
~#3 of 12 (estimated)~84.0%
No published score — estimated from related benchmarks
OSWorld 2.0Estimated
~#3 of 5 (estimated)~55.3%
No published score — estimated from related benchmarks
AutomationBench-AAVerified
#17 of 7956.6%
- #1Claude Opus 5.569.5%
- #2DeepSeek V4.1 Flash68.9%
- #16DeepSeek V4 Pro (0813)56.7%
- #17Claude Opus 556.6%
- #18Qwen3.8 Max (0902)56.2%
artificialanalysis.ai · 2026-09-24
Not comparable across labs (1)
Kept for reference, not counted in rankings: these use a lab-specific task set, answer key or scoring, so scores can't be compared fairly across labs. Only published scores are shown.