o3
- Input
- $2.00/M tok
- Output
- $8.00/M tok
- Blended
- $3.50/M tok · 3:1
- Speed
- 143tok/s
- Context
- 200Ktokens
Long-tail import from Artificial Analysis (2026-09-24); AA variant: o3. Price = AA median host price.
Rankings
Writing & chat
Verified#36 / 146rank by quality64.4±9.5quality / 1001/5 benchmarks measuredVision
Verified#37 / 82rank by quality56.0±8.2quality / 1001/6 benchmarks measuredHard reasoning
Verified#55 / 182rank by quality55.4±5.8quality / 1002/4 benchmarks measuredAgents
Verified#31 / 173rank by quality54.9±10.0quality / 1001/7 benchmarks measured
Benchmark scores
~ italic, dashed = no published score yet, estimated from related benchmarks.
Writing & chat
IFBenchVerified
#24 of 12571.4%
- #1MiniMax-M382.9%
- #2Qwen3.8 2.4T A95B82.8%
- #23Nemotron 3 Super 120B A12B71.5%
- #24o371.4%
- #25GPT-5.6 Terra71.2%
artificialanalysis.ai · 2026-09-24
LMArena Text, Non-English (Elo)Estimated
~#32 of 54 (estimated)~1419 Elo
- #1Claude Opus 5.51501
- #2Claude Fable 5.1Few votes1496
- #31Gemma 4 26B A4BFew votes1423
- ~#32o3Estimated1419
- #32Muse GlimmerFew votes1418
No published score — estimated from related benchmarks
LMArena Text (Elo)Estimated
~#33 of 54 (estimated)~1433 Elo
- #1Claude Opus 5.51509
- #2Claude Fable 5.11501
- #32MiMo-V2.51434
- ~#33o3Estimated1433
- #33Mistral Medium 3.51426
No published score — estimated from related benchmarks
EQ-Bench Creative Writing v3 (Elo)Estimated
~#13 of 26 (estimated)~1829 Elo
- #1GPT-6 Astra2173
- #2Claude Fable 5.12162
- #12Qwen3.8 2.4T A95B1843
- ~#13o3Estimated1829
- #13Muse Glimmer1798
No published score — estimated from related benchmarks
Vision
CharXiv ReasoningEstimated
~#7 of 9 (estimated)~77.5%
No published score — estimated from related benchmarks
OmniDocBench v1.5Estimated
~#5 of 8 (estimated)~84.6%
No published score — estimated from related benchmarks
GDP.pdf (AA)Estimated
~#26 of 43 (estimated)~12.6%
No published score — estimated from related benchmarks
ChartographyEstimated
~#15 of 23 (estimated)~18.2%
No published score — estimated from related benchmarks
Hard reasoning
GPQA DiamondVerified
#53 of 17282.7%
- #1GPT-6 Astra96.1%
- #2Gemini 3.8 Flash95.3%
- #52K-EXAONE 2.0 080382.9%
- #53o382.7%
- #54Qwen3.5 Omni Plus82.6%
artificialanalysis.ai · 2026-09-24
Humanity's Last ExamVerified
#62 of 17820.1%
- #1Claude Opus 5.561.4%
- #2Claude Fable 5.159.1%
- #61Nemotron 3 Super 120B A12B20.8%
- #62o320.1%
- #63GPT-5.5 Instant (June 2026)19.9%
artificialanalysis.ai · 2026-09-24
AIME (latest)Estimated
~#5 of 25 (estimated)~91.2%
No published score — estimated from related benchmarks
AA Intelligence IndexEstimated
~#36 of 75 (estimated)~25.1
No published score — estimated from related benchmarks
Agents
OSWorld-VerifiedEstimated
~#8 of 12 (estimated)~74.7%
No published score — estimated from related benchmarks
OSWorld 2.0Estimated
~#3 of 5 (estimated)~53.8%
No published score — estimated from related benchmarks
BrowseCompEstimated
~#15 of 18 (estimated)~68.8%
- #1Atria Dawn Preview92.5%
- #2GPT-6 Astra91.5%
- #14Inkling77.1%
- ~#15o3Estimated68.8%
- #15Mistral Medium 3.548.6%
No published score — estimated from related benchmarks
GDPval-AA (Elo)Estimated
~#23 of 87 (estimated)~1391 Elo
No published score — estimated from related benchmarks
AutomationBench-AAEstimated
~#30 of 80 (estimated)~26.2%
- #1Claude Opus 5.569.5%
- #2DeepSeek V4.1 Flash68.9%
- #29GPT-5.4 mini26.2%
- ~#30o3Estimated26.2%
- #30K2 Horizon MoVA 36B A4B26.1%
No published score — estimated from related benchmarks
τ-Bench Banking (AA)Estimated
~#33 of 89 (estimated)~21.9%
No published score — estimated from related benchmarks