Muse Spark 1.3
- Input
- $1.25/M tok
- Output
- $4.25/M tok
- Blended
- $2.00/M tok · 3:1
- Speed
- 219tok/s
- Context
- 1Mtokens
Rankings
Coding
Verified#8 / 103rank by quality74.9±4.5quality / 1004/7 benchmarks measuredSpeed frontierWriting & chat
Verified#10 / 146rank by quality82.3±4.7quality / 1003/5 benchmarks measuredSpeed frontierVision
Verified#6 / 82rank by quality79.1±4.8quality / 1003/6 benchmarks measuredHard reasoning
Verified#6 / 182rank by quality86.9±3.7quality / 1003/4 benchmarks measuredPrice frontierSpeed frontierAgents
Verified#3 / 173rank by quality87.4±5.8quality / 1003/7 benchmarks measuredPrice frontierSpeed frontier
Benchmark scores
~ italic, dashed = no published score yet, estimated from related benchmarks.
Coding
SWE-bench VerifiedEstimated
~#2 of 14 (estimated)~88.2%
No published score — estimated from related benchmarks
SWE-bench Pro (Public)Estimated
~#5 of 19 (estimated)~70.5%
- #1Claude Opus 5.589.9%
- #2Claude Sonnet 5.581.3%
- #4Claude Opus 579.2%
- ~#5Muse Spark 1.3Estimated70.5%
- #5Intern-S2-397B68.5%
No published score — estimated from related benchmarks
DeepSWE v1.1Verified
#4 of 3071.7%
- #1DeepSeek V4.1 Flash74.2%
- #2Grok 4.772.6%
- #3MiMo-V2.6-Pro71.9%
- #4Muse Spark 1.371.7%
- #5Claude Sonnet 5.571.0%
artificialanalysis.ai · 2026-09-24
Terminal-Bench 4.0Verified
#11 of 8333.3%
- #1Claude Sonnet 5.563.6%
- #2Claude Opus 5.559.6%
- #10MiMo-V2.6-Pro34.8%
- #11Muse Spark 1.333.3%
- #11Step 5 Preview33.3%
artificialanalysis.ai · 2026-09-24
LiveCodeBenchEstimated
~#2 of 20 (estimated)~92.2%
No published score — estimated from related benchmarks
SciCodeVerified
#8 of 9558.8%
- #1Claude Opus 5.566.9%
- #2Claude Fable 5.163.1%
- #7Step 5 Preview58.9%
- #8Muse Spark 1.358.8%
- #9Gemini 3.1 Pro58.7%
artificialanalysis.ai · 2026-09-24
Writing & chat
LMArena Text, Non-English (Elo)Few votesVerified
#3 of 531487 Elo
- #1Claude Opus 5.51501
- #2Claude Fable 5.1Few votes1496
- #3Muse Spark 1.3Few votes1487
- #4Claude Opus 51483
lmarena.ai · 2026-09-25
#9 of 251906 Elo
eqbench.com · 2026-09-24
Vision
MMMU-ProVerified
#8 of 7682.0%
- #1Claude Opus 5.587.7%
- #2GPT-6 Astra86.9%
- #7Gemini 3.1 Pro82.4%
- #8Muse Spark 1.382.0%
- #9Intern-S2-397B81.7%
artificialanalysis.ai · 2026-09-24
CharXiv ReasoningEstimated
~#4 of 9 (estimated)~84.1%
- #1Gemini 3.8 Flash86.2%
- #2Kimi K384.8%
- #3Qwen3.8 Flash-Next84.6%
- ~#4Muse Spark 1.3Estimated84.1%
- #4Qwen3.8-Omni-Flash83.5%
No published score — estimated from related benchmarks
OmniDocBench v1.5Estimated
~#4 of 8 (estimated)~90.5%
- #1MiniMax-M391.6%
- #2Kimi K391.1%
- #2Qwen3.8 27B91.1%
- ~#4Muse Spark 1.3Estimated90.5%
- #4Gemini 3.1 Pro85.3%
No published score — estimated from related benchmarks
ChartographyVerified
#10 of 2227.6%
- #1GPT-6 Astra71.0%
- #2Claude Opus 5.566.3%
- #8Qwen3.8 Max (0902)29.1%
- #10Muse Spark 1.327.6%
- #11Claude Opus 527.3%
surgehq.ai · 2026-09-24
Hard reasoning
GPQA DiamondVerified
#5 of 17293.5%
- #1GPT-6 Astra96.1%
- #2Gemini 3.8 Flash95.3%
- #5Kimi K393.5%
- #5Muse Spark 1.393.5%
- #5Qwen3.8 2.4T A95B93.5%
artificialanalysis.ai · 2026-09-24
Humanity's Last ExamVerified
#7 of 17848.7%
- #1Claude Opus 5.561.4%
- #2Claude Fable 5.159.1%
- #6MiMo-V2.6-Pro49.4%
- #7Muse Spark 1.348.7%
- #8GPT-6 Sol47.9%
artificialanalysis.ai · 2026-09-24
AIME (latest)Estimated
~#2 of 25 (estimated)~96.8%
No published score — estimated from related benchmarks
Agents
τ²-bench (Telecom)Estimated
~#2 of 116 (estimated)~98.3%
No published score — estimated from related benchmarks
OSWorld-VerifiedEstimated
~#3 of 12 (estimated)~84.2%
No published score — estimated from related benchmarks
OSWorld 2.0Estimated
~#3 of 5 (estimated)~57.8%
No published score — estimated from related benchmarks
BrowseCompEstimated
~#3 of 18 (estimated)~91.4%
No published score — estimated from related benchmarks
AutomationBench-AAVerified
#14 of 7957.9%
- #1Claude Opus 5.569.5%
- #2DeepSeek V4.1 Flash68.9%
- #13Kimi K358.3%
- #14Muse Spark 1.357.9%
- #15Qwen3.8 2.4T A95B57.2%
artificialanalysis.ai · 2026-09-24
Not comparable across labs (1)
Kept for reference, not counted in rankings: these use a lab-specific task set, answer key or scoring, so scores can't be compared fairly across labs. Only published scores are shown.