Gemini 3.1 Pro
- Input
- $2.00/M tok
- Output
- $12/M tok
- Blended
- $4.50/M tok · 3:1
- Speed
- 120tok/s
- Context
- 1Mtokens
Still in preview (gemini-3.1-pro-preview); Google's latest Pro model as of Sept 2026.
Rankings
Coding
Mixed#31 / 103rank by quality51.7±5.4quality / 1004/7 benchmarks measuredWriting & chat
Mixed#18 / 146rank by quality75.4±3.4quality / 1005/5 benchmarks measuredVision
Verified#19 / 82rank by quality68.6±4.0quality / 1004/6 benchmarks measuredHard reasoning
Verified#25 / 182rank by quality75.8±3.7quality / 1003/4 benchmarks measuredAgents
Mixed#32 / 173rank by quality54.7±3.4quality / 1005/7 benchmarks measured
Benchmark scores
~ italic, dashed = no published score yet, estimated from related benchmarks.
Coding
SWE-bench Pro (Public)Lab-reported
#15 of 1854.2%
- #1Claude Opus 5.589.9%
- #2Claude Sonnet 5.581.3%
- #14Inkling54.3%
- #15Gemini 3.1 Pro54.2%
- #15Gemini 3.5 Flash-Lite54.2%
deepmind.google · 2026-02-19
DeepSWE v1.1Estimated
~#22 of 31 (estimated)~59.8%
- #1DeepSeek V4.1 Flash74.2%
- #2Grok 4.772.6%
- #21GLM-5.361.4%
- ~#22Gemini 3.1 ProEstimated59.8%
- #22Qwen3.8 Flash-Next58.7%
No published score — estimated from related benchmarks
Terminal-Bench 4.0Verified
#28 of 834.0%
- #1Claude Sonnet 5.563.6%
- #2Claude Opus 5.559.6%
- #27Qwen3.8 27B5.6%
- #28Gemini 3.1 Pro4.0%
- #29GPT-5.4 mini2.0%
artificialanalysis.ai · 2026-09-24
LiveCodeBenchEstimated
~#6 of 20 (estimated)~85.6%
- #1Qwen3.8-Omni-Flash92.6%
- #2Qwen3.8 Flash-Next91.9%
- #5gpt-oss-120b87.8%
- ~#6Gemini 3.1 ProEstimated85.6%
- #6Gemma 4 31B80.0%
No published score — estimated from related benchmarks
SciCodeVerified
#9 of 9558.7%
- #1Claude Opus 5.566.9%
- #2Claude Fable 5.163.1%
- #8Muse Spark 1.358.8%
- #9Gemini 3.1 Pro58.7%
- #10GPT-6 Sol57.6%
artificialanalysis.ai · 2026-09-24
AA Coding Agent IndexEstimated
~#11 of 15 (estimated)~46.6
- #1Claude Opus 5.566.0
- #2Claude Fable 5.162.2
- #10Grok 4.647.0
- ~#11Gemini 3.1 ProEstimated46.6
- #11Qwen3.8 Max (0902)43.3
No published score — estimated from related benchmarks
Writing & chat
IFBenchVerified
#13 of 12577.1%
- #1MiniMax-M382.9%
- #2Qwen3.8 2.4T A95B82.8%
- #12Qwen3.7 Plus78.0%
- #13Gemini 3.1 Pro77.1%
- #14Muse Glimmer77.0%
artificialanalysis.ai · 2026-09-24
LMArena Text, Non-English (Elo)Verified
#6 of 531480 Elo
- #1Claude Opus 5.51501
- #2Claude Fable 5.1Few votes1496
- #4Gemini 3.8 FlashFew votes1483
- #6Gemini 3.1 Pro1480
- #7Kimi K31477
lmarena.ai · 2026-09-24
Vision
MMMU-ProVerified
#7 of 7682.4%
- #1Claude Opus 5.587.7%
- #2GPT-6 Astra86.9%
- #6Qwen3.8 Max (0902)82.8%
- #7Gemini 3.1 Pro82.4%
- #8Muse Spark 1.382.0%
artificialanalysis.ai · 2026-09-24
CharXiv ReasoningEstimated
~#5 of 9 (estimated)~82.0%
- #1Gemini 3.8 Flash86.2%
- #2Kimi K384.8%
- #4Qwen3.8-Omni-Flash83.5%
- ~#5Gemini 3.1 ProEstimated82.0%
- #5Muse Glimmer78.8%
No published score — estimated from related benchmarks
Hard reasoning
Humanity's Last ExamVerified
#10 of 17847.0%
- #1Claude Opus 5.561.4%
- #2Claude Fable 5.159.1%
- #9Gemini 3.8 Flash47.8%
- #10Gemini 3.1 Pro47.0%
- #11Kimi K346.9%
artificialanalysis.ai · 2026-09-24
AIME (latest)Estimated
~#2 of 25 (estimated)~96.7%
No published score — estimated from related benchmarks
Agents
OSWorld-VerifiedEstimated
~#9 of 12 (estimated)~73.9%
- #1Kimi K384.8%
- #2Qwen3.8 27B84.3%
- #8Gemini 3.5 Flash-Lite74.0%
- ~#9Gemini 3.1 ProEstimated73.9%
- #9Nex-N2.5-Mini71.2%
No published score — estimated from related benchmarks
OSWorld 2.0Estimated
~last of 5 (estimated)below measured range (<47.5%)
No published score — estimated from related benchmarks
BrowseCompLab-reported
#8 of 1785.9%
- #1Atria Dawn Preview92.5%
- #2GPT-6 Astra91.5%
- #7GPT-5.6 Terra87.5%
- #8Gemini 3.1 Pro85.9%
- #9Claude Sonnet 584.7%
deepmind.google · 2026-02-19
GDPval-AA (Elo)Verified
#48 of 86776 Elo
- #1Claude Opus 5.51846
- #2Claude Fable 5.11735
- #47Qwen3.5 397B A17B780
- #48Gemini 3.1 Pro776
- #49Muse Glimmer774
artificialanalysis.ai · 2026-09-24