Grok 4.6
- Input
- $2.00/M tok
- Output
- $6.00/M tok
- Blended
- $3.00/M tok · 3:1
- Speed
- 60tok/s
- Context
- 500Ktokens
Previous generation.
Rankings
Coding
Verified#25 / 103rank by quality59.0±4.5quality / 1004/7 benchmarks measuredWriting & chat
Verified#22 / 146rank by quality72.0±6.0quality / 1002/5 benchmarks measuredVision
Verified#21 / 82rank by quality65.2±7.3quality / 1002/6 benchmarks measuredHard reasoning
Verified#12 / 182rank by quality82.2±3.7quality / 1003/4 benchmarks measuredAgents
Verified#10 / 173rank by quality84.1±5.8quality / 1003/7 benchmarks measured
Benchmark scores
~ italic, dashed = no published score yet, estimated from related benchmarks.
Coding
SWE-bench VerifiedEstimated
~#3 of 14 (estimated)~82.8%
No published score — estimated from related benchmarks
SWE-bench Pro (Public)Estimated
~#7 of 19 (estimated)~64.1%
- #1Claude Opus 5.589.9%
- #2Claude Sonnet 5.581.3%
- #6Hy4 preview65.7%
- ~#7Grok 4.6Estimated64.1%
- #7GPT-5.6 Terra63.4%
No published score — estimated from related benchmarks
DeepSWE v1.1Verified
#15 of 3064.9%
- #1DeepSeek V4.1 Flash74.2%
- #2Grok 4.772.6%
- #14Gemini 3.8 Flash65.8%
- #15Grok 4.664.9%
- #16Claude Fable 5.164.3%
artificialanalysis.ai · 2026-09-24
Terminal-Bench 4.0Verified
#19 of 8317.2%
- #1Claude Sonnet 5.563.6%
- #2Claude Opus 5.559.6%
- #18Gemini 3.8 Flash19.7%
- #19Grok 4.617.2%
- #20GPT-5.514.6%
artificialanalysis.ai · 2026-09-24
LiveCodeBenchEstimated
~#3 of 20 (estimated)~90.8%
No published score — estimated from related benchmarks
Writing & chat
IFBenchEstimated
~#20 of 126 (estimated)~74.7%
- #1MiniMax-M382.9%
- #2Qwen3.8 2.4T A95B82.8%
- #19GPT-5.3 Codex75.4%
- ~#20Grok 4.6Estimated74.7%
- #20Command A+73.9%
No published score — estimated from related benchmarks
LMArena Text, Non-English (Elo)Verified
#18 of 531447 Elo
- #1Claude Opus 5.51501
- #2Claude Fable 5.1Few votes1496
- #18Claude Sonnet 51447
- #18Grok 4.61447
- #20Gemini 3.5 Flash-Lite1445
lmarena.ai · 2026-09-24
EQ-Bench Creative Writing v3 (Elo)Estimated
~#10 of 26 (estimated)~1897 Elo
- #1GPT-6 Astra2173
- #2Claude Fable 5.12162
- #9Muse Spark 1.31906
- ~#10Grok 4.6Estimated1897
- #10GPT-5.6 Terra1855
No published score — estimated from related benchmarks
Vision
MMMU-ProEstimated
~#10 of 77 (estimated)~81.5%
- #1Claude Opus 5.587.7%
- #2GPT-6 Astra86.9%
- #9Intern-S2-397B81.7%
- ~#10Grok 4.6Estimated81.5%
- #10GPT-5.6 Terra80.7%
No published score — estimated from related benchmarks
CharXiv ReasoningEstimated
~#5 of 9 (estimated)~81.8%
- #1Gemini 3.8 Flash86.2%
- #2Kimi K384.8%
- #4Qwen3.8-Omni-Flash83.5%
- ~#5Grok 4.6Estimated81.8%
- #5Muse Glimmer78.8%
No published score — estimated from related benchmarks
OmniDocBench v1.5Estimated
~#4 of 8 (estimated)~88.2%
No published score — estimated from related benchmarks
Hard reasoning
AIME (latest)Estimated
~#2 of 25 (estimated)~96.8%
No published score — estimated from related benchmarks
Agents
τ²-bench (Telecom)Estimated
~#2 of 116 (estimated)~97.7%
No published score — estimated from related benchmarks
OSWorld-VerifiedEstimated
~#3 of 12 (estimated)~82.4%
No published score — estimated from related benchmarks
OSWorld 2.0Estimated
~#3 of 5 (estimated)~56.6%
No published score — estimated from related benchmarks
BrowseCompEstimated
~#6 of 18 (estimated)~89.6%
- #1Atria Dawn Preview92.5%
- #2GPT-6 Astra91.5%
- #5Nex-N2.5-Pro89.7%
- ~#6Grok 4.6Estimated89.6%
- #6Step 5 Preview88.7%
No published score — estimated from related benchmarks
GDPval-AA (Elo)Verified
#10 of 861632 Elo
- #1Claude Opus 5.51846
- #2Claude Fable 5.11735
- #9GLM-5.3 Flash1641
- #10Grok 4.61632
- #11Qwen3.8 Flash-Next1612
artificialanalysis.ai · 2026-09-24