GPT-6 Astra
- Input
- $10/M tok
- Output
- $50/M tok
- Blended
- $20/M tok · 3:1
- Speed
- 52tok/s
- Context
- 1Mtokens
Rankings
Coding
Verified#4 / 103rank by quality86.2±4.5quality / 1004/7 benchmarks measuredWriting & chat
Verified#4 / 146rank by quality92.1±5.3quality / 1003/5 benchmarks measuredVision
Verified#1 / 82rank by quality98.9±3.0quality / 1003/6 benchmarks measuredPrice frontierSpeed frontierHard reasoning
Verified#4 / 182rank by quality93.5±3.7quality / 1003/4 benchmarks measuredAgents
Mixed#15 / 173rank by quality81.6±4.6quality / 1004/7 benchmarks measured
Benchmark scores
~ italic, dashed = no published score yet, estimated from related benchmarks.
Coding
SWE-bench VerifiedEstimated
~#2 of 14 (estimated)~90.0%
No published score — estimated from related benchmarks
SWE-bench Pro (Public)Estimated
~#5 of 19 (estimated)~75.1%
- #1Claude Opus 5.589.9%
- #2Claude Sonnet 5.581.3%
- #4Claude Opus 579.2%
- ~#5GPT-6 AstraEstimated75.1%
- #5Intern-S2-397B68.5%
No published score — estimated from related benchmarks
LiveCodeBenchEstimated
~#2 of 20 (estimated)~92.3%
No published score — estimated from related benchmarks
SciCodeVerified
#13 of 9556.5%
- #1Claude Opus 5.566.9%
- #2Claude Fable 5.163.1%
- #12Gemini 3.8 Flash56.6%
- #13GPT-6 Astra56.5%
- #14Claude Opus 556.4%
artificialanalysis.ai · 2026-09-24
Writing & chat
LMArena Text, Non-English (Elo)Few votesVerified
#14 of 531460 Elo
- #1Claude Opus 5.51501
- #2Claude Fable 5.1Few votes1496
- #13MiMo-V2.6-Pro1461
- #14GPT-6 AstraFew votes1460
- #15GPT-5.6 Terra1456
lmarena.ai · 2026-09-24
LMArena Text (Elo)Few votesVerified
#10 of 531480 Elo
- #1Claude Opus 5.51509
- #2Claude Fable 5.11501
- #10GLM-5.31480
- #10GPT-6 AstraFew votes1480
- #10MiMo-V2.6-Pro1480
lmarena.ai · 2026-09-24
#1 of 252173 Elo
eqbench.com · 2026-09-24
Vision
CharXiv ReasoningEstimated
~#2 of 9 (estimated)~86.1%
No published score — estimated from related benchmarks
OmniDocBench v1.5Estimated
~#2 of 8 (estimated)~91.4%
No published score — estimated from related benchmarks
Hard reasoning
Humanity's Last ExamVerified
#5 of 17854.7%
- #1Claude Opus 5.561.4%
- #2Claude Fable 5.159.1%
- #4Claude Opus 554.9%
- #5GPT-6 Astra54.7%
- #6MiMo-V2.6-Pro49.4%
artificialanalysis.ai · 2026-09-24
AIME (latest)Estimated
~#2 of 25 (estimated)~96.9%
No published score — estimated from related benchmarks
Agents
τ²-bench (Telecom)Estimated
~#2 of 116 (estimated)~98.3%
No published score — estimated from related benchmarks
OSWorld-VerifiedEstimated
~#3 of 12 (estimated)~83.4%
No published score — estimated from related benchmarks
OSWorld 2.0Estimated
~#3 of 5 (estimated)~55.0%
No published score — estimated from related benchmarks
Not comparable across labs (1)
Kept for reference, not counted in rankings: these use a lab-specific task set, answer key or scoring, so scores can't be compared fairly across labs. Only published scores are shown.