GPT-6 Luna
- Input
- $0.10/M tok
- Output
- $0.50/M tok
- Blended
- $0.20/M tok · 3:1
- Speed
- 132tok/s
- Context
- 1Mtokens
Writing: no LMArena/EQ-Bench entry or lab-reported writing benchmark as of 2026-09-24, so its Writing entry is fully estimated from its other categories (flagged “Estimated from other categories”).
Rankings
Coding
Verified#30 / 103rank by quality52.5±4.5quality / 1004/7 benchmarks measuredWriting & chat
Verified#33 / 146rank by quality67.2±6.0quality / 1002/5 benchmarks measuredPrice frontierVision
Verified#18 / 82rank by quality68.8±4.8quality / 1003/6 benchmarks measuredPrice frontierHard reasoning
Verified#28 / 182rank by quality72.4±5.1quality / 1002/4 benchmarks measuredPrice frontierAgents
Verified#29 / 173rank by quality64.2±7.8quality / 1002/7 benchmarks measured
Benchmark scores
~ italic, dashed = no published score yet, estimated from related benchmarks.
Coding
SWE-bench VerifiedEstimated
~#3 of 14 (estimated)~82.0%
No published score — estimated from related benchmarks
SWE-bench Pro (Public)Estimated
~#9 of 19 (estimated)~62.5%
- #1Claude Opus 5.589.9%
- #2Claude Sonnet 5.581.3%
- #8Claude Sonnet 563.2%
- ~#9GPT-6 LunaEstimated62.5%
- #9Nex-N2.5-Pro61.2%
No published score — estimated from related benchmarks
Terminal-Bench 4.0Verified
#23 of 8312.6%
- #1Claude Sonnet 5.563.6%
- #2Claude Opus 5.559.6%
- #23GPT-5.5 Instant (June 2026)12.6%
- #23GPT-6 Luna12.6%
- #23Kimi K312.6%
artificialanalysis.ai · 2026-09-24
LiveCodeBenchEstimated
~#6 of 20 (estimated)~85.4%
- #1Qwen3.8-Omni-Flash92.6%
- #2Qwen3.8 Flash-Next91.9%
- #5gpt-oss-120b87.8%
- ~#6GPT-6 LunaEstimated85.4%
- #6Gemma 4 31B80.0%
No published score — estimated from related benchmarks
SciCodeVerified
#17 of 9554.6%
- #1Claude Opus 5.566.9%
- #2Claude Fable 5.163.1%
- #16GPT-5.6 Terra55.0%
- #17GPT-6 Luna54.6%
- #18Claude Sonnet 554.3%
artificialanalysis.ai · 2026-09-24
Writing & chat
IFBenchEstimated
~#23 of 126 (estimated)~71.8%
- #1MiniMax-M382.9%
- #2Qwen3.8 2.4T A95B82.8%
- #22Gemma 4 26B A4B72.4%
- ~#23GPT-6 LunaEstimated71.8%
- #23Nemotron 3 Super 120B A12B71.5%
No published score — estimated from related benchmarks
LMArena Text, Non-English (Elo)Verified
#25 of 531436 Elo
- #1Claude Opus 5.51501
- #2Claude Fable 5.1Few votes1496
- #24Gemma 4 31BFew votes1437
- #25GPT-6 Luna1436
- #26Grok 4.71435
lmarena.ai · 2026-09-25
LMArena Text (Elo)Verified
#25 of 531442 Elo
- #1Claude Opus 5.51509
- #2Claude Fable 5.11501
- #24GPT-5.4 mini1448
- #25GPT-6 Luna1442
- #25Qwen3.5 397B A17B1442
lmarena.ai · 2026-09-25
EQ-Bench Creative Writing v3 (Elo)Estimated
~#10 of 26 (estimated)~1862 Elo
- #1GPT-6 Astra2173
- #2Claude Fable 5.12162
- #9Muse Spark 1.31906
- ~#10GPT-6 LunaEstimated1862
- #10GPT-5.6 Terra1855
No published score — estimated from related benchmarks
Vision
CharXiv ReasoningEstimated
~#5 of 9 (estimated)~81.5%
- #1Gemini 3.8 Flash86.2%
- #2Kimi K384.8%
- #4Qwen3.8-Omni-Flash83.5%
- ~#5GPT-6 LunaEstimated81.5%
- #5Muse Glimmer78.8%
No published score — estimated from related benchmarks
OmniDocBench v1.5Estimated
~#4 of 8 (estimated)~87.3%
No published score — estimated from related benchmarks
Hard reasoning
GPQA DiamondEstimated
~#33 of 173 (estimated)~87.9%
- #1GPT-6 Astra96.1%
- #2Gemini 3.8 Flash95.3%
- #32Solar Pro 489.1%
- ~#33GPT-6 LunaEstimated87.9%
- #33GPT-5.4 mini87.5%
No published score — estimated from related benchmarks
Humanity's Last ExamVerified
#27 of 17838.5%
- #1Claude Opus 5.561.4%
- #2Claude Fable 5.159.1%
- #26MiniMax-M339.0%
- #27GPT-6 Luna38.5%
- #28Grok Build 0.1 061638.3%
artificialanalysis.ai · 2026-09-24
AIME (latest)Estimated
~#3 of 25 (estimated)~94.4%
No published score — estimated from related benchmarks
Agents
τ²-bench (Telecom)Estimated
~#16 of 116 (estimated)~89.2%
- #1Step 3.7 Flash98.5%
- #2Gemini 3.1 Pro95.6%
- #15KAT Coder Pro V289.5%
- ~#16GPT-6 LunaEstimated89.2%
- #16MiniMax-M388.9%
No published score — estimated from related benchmarks
OSWorld-VerifiedEstimated
~#8 of 12 (estimated)~78.7%
- #1Kimi K384.8%
- #2Qwen3.8 27B84.3%
- #7GPT-5.578.7%
- ~#8GPT-6 LunaEstimated78.7%
- #8Gemini 3.5 Flash-Lite74.0%
No published score — estimated from related benchmarks
OSWorld 2.0Estimated
~#3 of 5 (estimated)~50.4%
No published score — estimated from related benchmarks
BrowseCompEstimated
~#14 of 18 (estimated)~81.8%
- #1Atria Dawn Preview92.5%
- #2GPT-6 Astra91.5%
- #13Nex-N2.5-Mini83.4%
- ~#14GPT-6 LunaEstimated81.8%
- #14Inkling77.1%
No published score — estimated from related benchmarks
GDPval-AA (Elo)Verified
#23 of 861367 Elo
- #1Claude Opus 5.51846
- #2Claude Fable 5.11735
- #22Qwen3.8 27B1409
- #23GPT-6 Luna1367
- #24K2 Horizon 375B A23B1349
artificialanalysis.ai · 2026-09-24
AutomationBench-AAVerified
#20 of 7953.2%
- #1Claude Opus 5.569.5%
- #2DeepSeek V4.1 Flash68.9%
- #19Qwen3.8 Flash-Next55.9%
- #20GPT-6 Luna53.2%
- #21Step 5 Preview51.0%
artificialanalysis.ai · 2026-09-24
τ-Bench Banking (AA)Estimated
~#22 of 89 (estimated)~31.5%
- #1Muse Spark 1.350.5%
- #2GLM-5.350.3%
- #21K2 Horizon MoVA 36B A4B32.0%
- ~#22GPT-6 LunaEstimated31.5%
- #22Inkling29.1%
No published score — estimated from related benchmarks
Not comparable across labs (1)
Kept for reference, not counted in rankings: these use a lab-specific task set, answer key or scoring, so scores can't be compared fairly across labs. Only published scores are shown.