GPT-5.5
- Input
- $5.00/M tok
- Output
- $30/M tok
- Blended
- $11.3/M tok · 3:1
- Speed
- 91tok/s
- Context
- 922Ktokens
Previous generation.
Rankings
Coding
Mixed#22 / 103rank by quality60.2±5.2quality / 1004/7 benchmarks measuredWriting & chat
Verified#16 / 146rank by quality77.0±3.3quality / 1004/5 benchmarks measuredVision
Verified#14 / 82rank by quality74.4±4.8quality / 1003/6 benchmarks measuredHard reasoning
Verified#17 / 182rank by quality79.9±3.7quality / 1003/4 benchmarks measuredAgents
Mixed#26 / 173rank by quality67.4±3.2quality / 1007/7 benchmarks measured
Benchmark scores
~ italic, dashed = no published score yet, estimated from related benchmarks.
Coding
SWE-bench VerifiedEstimated
~#3 of 14 (estimated)~82.9%
No published score — estimated from related benchmarks
Terminal-Bench 4.0Verified
#20 of 8314.6%
- #1Claude Sonnet 5.563.6%
- #2Claude Opus 5.559.6%
- #19Grok 4.617.2%
- #20GPT-5.514.6%
- #21Claude Sonnet 514.1%
artificialanalysis.ai · 2026-09-24
LiveCodeBenchEstimated
~#3 of 20 (estimated)~90.6%
No published score — estimated from related benchmarks
SciCodeVerified
#15 of 9555.8%
- #1Claude Opus 5.566.9%
- #2Claude Fable 5.163.1%
- #14Claude Opus 556.4%
- #15GPT-5.555.8%
- #16GPT-5.6 Terra55.0%
artificialanalysis.ai · 2026-09-24
AA Coding Agent IndexEstimated
~#10 of 15 (estimated)~48.9
No published score — estimated from related benchmarks
Writing & chat
IFBenchVerified
#15 of 12575.9%
- #1MiniMax-M382.9%
- #2Qwen3.8 2.4T A95B82.8%
- #15GPT-5.4 nano75.9%
- #15GPT-5.575.9%
- #17Qwen3.5 122B A10B75.7%
artificialanalysis.ai · 2026-09-24
LMArena Text, Non-English (Elo)Verified
#8 of 531474 Elo
- #1Claude Opus 5.51501
- #2Claude Fable 5.1Few votes1496
- #7Kimi K31477
- #8GPT-5.51474
- #9Qwen3.8 Max (0902)1467
lmarena.ai · 2026-09-24
LMArena Text (Elo)Verified
#8 of 531482 Elo
- #1Claude Opus 5.51509
- #2Claude Fable 5.11501
- #7Gemini 3.1 Pro1487
- #8GPT-5.51482
- #9Qwen3.8 Max (0902)1481
lmarena.ai · 2026-09-24
#11 of 251844 Elo
eqbench.com · 2026-09-24
Vision
MMMU-ProVerified
#13 of 7679.9%
- #1Claude Opus 5.587.7%
- #2GPT-6 Astra86.9%
- #11Qwen3.7 Plus80.5%
- #13GPT-5.579.9%
- #14Qwen3.8 Flash-Next79.8%
artificialanalysis.ai · 2026-09-24
CharXiv ReasoningEstimated
~#5 of 9 (estimated)~83.3%
- #1Gemini 3.8 Flash86.2%
- #2Kimi K384.8%
- #4Qwen3.8-Omni-Flash83.5%
- ~#5GPT-5.5Estimated83.3%
- #5Muse Glimmer78.8%
No published score — estimated from related benchmarks
OmniDocBench v1.5Estimated
~#4 of 8 (estimated)~89.8%
No published score — estimated from related benchmarks
Hard reasoning
AIME (latest)Estimated
~#2 of 25 (estimated)~96.8%
No published score — estimated from related benchmarks
Agents
τ²-bench (Telecom)Verified
#6 of 11593.9%
- #1Step 3.7 Flash98.5%
- #2Gemini 3.1 Pro95.6%
- #5Mistral Medium 3.594.2%
- #6GPT-5.593.9%
- #7Qwen3.5 122B A10B93.6%
artificialanalysis.ai · 2026-09-24
GDPval-AA (Elo)Verified
#25 of 861336 Elo
- #1Claude Opus 5.51846
- #2Claude Fable 5.11735
- #24K2 Horizon 375B A23B1349
- #25GPT-5.51336
- #26MiniMax-M31230
artificialanalysis.ai · 2026-09-24