Mistral Large 4
- Input
- $1.36/M tok
- Output
- $4.18/M tok
- Blended
- $2.065/M tok · 3:1
- Speed
- ~100tok/sest.
- Context
- —
- Price
- —
- Time
- —
Preview; weights "drop end of this month" and are not downloadable yet. Launch post states 1T total / 49B active; once weights ship, set openWeights=true, re-add params and verify against the HF card.
Rankings
Benchmark scores
~ italic, dashed = no published score yet, estimated from related benchmarks.
Coding
SWE-bench VerifiedEstimated
~#2 of 14 (estimated)~87.0%
No published score — estimated from related benchmarks
SWE-bench Pro (Public)Estimated
~#6 of 20 (estimated)~68.0%
- #1Claude Opus 5.589.9%
- #2Claude Sonnet 5.581.3%
- #5Intern-S2-397B68.5%
- ~#6Mistral Large 4Estimated68.0%
- #6Hy4 preview65.7%
No published score — estimated from related benchmarks
DeepSWE v1.1Lab-reported
#23 of 3461.7%
- #1Gemini 4 Argon78.8%
- #2DeepSeek V4.1 Flash74.2%
- #22Claude Opus 562.5%
- #23Mistral Large 461.7%
- #24GLM-5.361.4%
mistral.ai · 2026-10-07
Terminal-Bench 4.0Lab-reported
#17 of 8828.3%
- #1Claude Sonnet 5.563.6%
- #2Claude Opus 5.559.6%
- #15GLM-5.3 Flash32.8%
- #17Mistral Large 428.3%
- #18DeepSeek V4.1 Flash26.8%
mistral.ai · 2026-10-07
LiveCodeBenchEstimated
~#6 of 25 (estimated)~88.5%
- #1Qwen3.8-Omni-Flash92.6%
- #2Qwen3.8 Flash-Next91.9%
- #5Nemotron 3 Ultra89.0%
- ~#6Mistral Large 4Estimated88.5%
- #6gpt-oss-120b87.8%
No published score — estimated from related benchmarks
SciCodeEstimated
~#27 of 106 (estimated)~53.3%
- #1Claude Opus 5.566.9%
- #2Claude Fable 5.163.1%
- #26DeepSeek-V4-Pro-0813-nvfp4-DSpark53.8%
- ~#27Mistral Large 4Estimated53.3%
- #27Grok 4.653.0%
No published score — estimated from related benchmarks
Agents
τ²-bench (Telecom)Estimated
~#12 of 118 (estimated)~91.7%
- #1Step 3.7 Flash98.5%
- #2DeepSeek-V4-Pro-0813-nvfp4-DSpark98.3%
- #11Ring-2.6-1T92.4%
- ~#12Mistral Large 4Estimated91.7%
- #12MiMo-V2.590.6%
No published score — estimated from related benchmarks
OSWorld-VerifiedEstimated
~#3 of 12 (estimated)~83.4%
No published score — estimated from related benchmarks
OSWorld 2.0Estimated
~#3 of 5 (estimated)~55.6%
No published score — estimated from related benchmarks
BrowseCompEstimated
~#5 of 18 (estimated)~90.5%
- #1Atria Dawn Preview92.5%
- #2GPT-6 Astra91.5%
- #4Claude Opus 590.8%
- ~#5Mistral Large 4Estimated90.5%
- #5Nex-N2.5-Pro89.7%
No published score — estimated from related benchmarks
GDPval-AA (Elo)Estimated
~#14 of 88 (estimated)~1585 Elo
- #1Claude Opus 5.51866
- #2Claude Fable 5.11758
- #13Qwen3.8 2.4T A95B1597
- ~#14Mistral Large 4Estimated1585
- #14GPT-6.1 Sol1575
No published score — estimated from related benchmarks
AutomationBench-AALab-reported
#10 of 8159.9%
- #1Claude Opus 5.569.5%
- #2DeepSeek V4.1 Flash68.9%
- #10Gemini 3.8 Flash59.9%
- #10Mistral Large 459.9%
- #12GPT-5.6 Terra59.6%
mistral.ai · 2026-10-07
τ-Bench Banking (AA)Estimated
~#12 of 90 (estimated)~43.2%
- #1Muse Spark 1.350.5%
- #2GLM-5.350.3%
- #11Grok 4.643.3%
- ~#12Mistral Large 4Estimated43.2%
- #12Claude Opus 542.1%
No published score — estimated from related benchmarks