Skip to content
Model Pareto

Compare models

Side by side: quality per category, price, speed, and every benchmark cell with where it came from. Pick up to four.
GPT-6 LunaGrok Build 0.1 0616MiniMax-M3
Focus

Vision: quality vs price

Compared models are ringed; the other 79 ranked here are greyed.

Lower price is better. Pareto frontier: Gemma 4 E4B, Qwen3.5 4B, Ling-3.0-flash-VL, MiMo-V2.5, GPT-6 Luna, Qwen3.8-Omni-Flash, Gemini 3.8 Flash, GPT-6 Sol, Claude Opus 5.5, GPT-6 Astra. 16 models are not plotted: Intern-S2-397B, Intern-S2-Preview (35B-A3B), EXAONE 4.5 33B, DiffusionGemma 26B A4B, Step3 VL 10B, Command A+, Devstral Small 2, Qwen3.5 2B, Gemma 4 E2B, Llama 3.2 Instruct 90B (Vision), MiniCPM-V 4.6 1.3B, Molmo2-8B, LFM2.5-VL-1.6B, Qwen3.5 0.8B, Molmo 7B-D, Phi-4 Multimodal Instruct.

16 models with no price data — shown in the strip at the left edge

Intern-S2-397B, Intern-S2-Preview (35B-A3B), EXAONE 4.5 33B, DiffusionGemma 26B A4B, Step3 VL 10B, Command A+, Devstral Small 2, Qwen3.5 2B, Gemma 4 E2B, Llama 3.2 Instruct 90B (Vision), MiniCPM-V 4.6 1.3B, Molmo2-8B, LFM2.5-VL-1.6B, Qwen3.5 0.8B, Molmo 7B-D, Phi-4 Multimodal Instruct

  • Best-value frontier (nothing is both cheaper and better)
  • Evidencestrong → weak
  • Estimated from other categories
  • GPT-6 Luna
    Quality
    68.8
    Rank
    #18/82
    Price
    $0.20/M tok
    Speed
    132tok/s
    In$0.10▾Out$0.50▾/M tok

    vs MiniMax-M3: +5.0 quality · 0.38× price

  • Grok Build 0.1 0616
    Quality
    62.9
    Rank
    #26/82
    Price
    $1.25/M tok
    Speed
    55tok/s
    In$1.00Out$2.00/M tok

    vs GPT-6 Luna: −5.9 quality · 6.3× price

  • MiniMax-M3
    Quality
    63.8
    Rank
    #24/82
    Price
    $0.52/M tok
    Speed
    118tok/s
    In$0.30Out$1.20/M tok

    vs GPT-6 Luna: −5.0 quality · 2.6× price

Overview

Lab
OpenAI
SpaceXAI (xAI)
MiniMax
Released
Sep 22, 2026
Jun 16, 2026
Jun 1, 2026
Weights
Proprietary
Proprietary
Open
Input price
$/M tok
$0.10Best
$1.00
$0.30
Output price
$/M tok
$0.50Best
$2.00
$1.20
Blended price
$/M tok · 3:1 in:out
$0.20Best
$1.25
$0.525
Output speed
tok/s
132Best
55
118
⚡ SambaNova 246
Context
tokens
1MBest
256K
1MBest

Quality by category

0–100 within each category (100 = best tracked model). Rank is among all models ranked there. Click a row to focus it.

52.5±4.5Best
#30/103Verified
Not ranked
44.1±5.4
#37/103Mixed
67.2±6.0
#33/146Verifiedbest value
Not ranked
71.1±4.2Best
#24/146Verified
68.8±4.8Best
#18/82Verifiedbest value
62.9±8.2
#26/82Verified
63.8±5.3
#24/82Mixed
72.4±5.1Best
#28/182Verifiedbest value
72.0±5.8
#29/182Verified
70.7±3.7
#30/182Verified
64.2±7.8Best
#29/173Verified
38.6±7.8
#67/173Verified
47.4±3.4
#38/173Mixed

Vision benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

MMMU
day 0
~#2/3~71.3%Estimated
estimated from related benchmarks
~#2/3~71.3%Estimated
estimated from related benchmarks
~#4/8~87.3%Estimated
estimated from related benchmarks
#1/791.6%Lab-reported
minimax.io · 2026-06-01
Benchmarks in other categories (4)

Coding benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

LiveCodeBench
third-party
—

Writing & chat benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

Hard reasoning benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

~#3/25~94.4%Estimated
estimated from related benchmarks
~#2/25~95.8%Estimated
estimated from related benchmarks

Agents benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

~last/5below measured range (<47.5%)Estimated
estimated from related benchmarks
~last/5below measured range (<47.5%)Estimated
estimated from related benchmarks
Not comparable across labs (2)
—
—
#1/170.1%Lab-reported
minimax.io · 2026-06-01