Compare models
Vision: quality vs price
Compared models are ringed; the other 79 ranked here are greyed.
Lower price is better. Pareto frontier: Gemma 4 E4B, Qwen3.5 4B, Ling-3.0-flash-VL, MiMo-V2.5, GPT-6 Luna, Qwen3.8-Omni-Flash, Gemini 3.8 Flash, GPT-6 Sol, Claude Opus 5.5, GPT-6 Astra. 16 models are not plotted: Intern-S2-397B, Intern-S2-Preview (35B-A3B), EXAONE 4.5 33B, DiffusionGemma 26B A4B, Step3 VL 10B, Command A+, Devstral Small 2, Qwen3.5 2B, Gemma 4 E2B, Llama 3.2 Instruct 90B (Vision), MiniCPM-V 4.6 1.3B, Molmo2-8B, LFM2.5-VL-1.6B, Qwen3.5 0.8B, Molmo 7B-D, Phi-4 Multimodal Instruct.
16 models with no price data — shown in the strip at the left edge
Intern-S2-397B, Intern-S2-Preview (35B-A3B), EXAONE 4.5 33B, DiffusionGemma 26B A4B, Step3 VL 10B, Command A+, Devstral Small 2, Qwen3.5 2B, Gemma 4 E2B, Llama 3.2 Instruct 90B (Vision), MiniCPM-V 4.6 1.3B, Molmo2-8B, LFM2.5-VL-1.6B, Qwen3.5 0.8B, Molmo 7B-D, Phi-4 Multimodal Instruct
- Best-value frontier (nothing is both cheaper and better)
- Evidencestrong → weak
- Estimated from other categories
- Hover a dot for details · click a lab to highlight it
- GPT-6 Luna
- Quality
- 68.8
- Rank
- #18/82
- Price
- $0.20/M tok
- Speed
- 132tok/s
In$0.10▾Out$0.50▾/M tokvs MiniMax-M3: +5.0 quality · 0.38× price
- Grok Build 0.1 0616
- Quality
- 62.9
- Rank
- #26/82
- Price
- $1.25/M tok
- Speed
- 55tok/s
In$1.00Out$2.00/M tokvs GPT-6 Luna: −5.9 quality · 6.3× price
- MiniMax-M3
- Quality
- 63.8
- Rank
- #24/82
- Price
- $0.52/M tok
- Speed
- 118tok/s
In$0.30Out$1.20/M tokvs GPT-6 Luna: −5.0 quality · 2.6× price
Overview
Quality by category
0–100 within each category (100 = best tracked model). Rank is among all models ranked there. Click a row to focus it.
Vision benchmarks
Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.
Benchmarks in other categories (4)· Coding, Writing & chat, Hard reasoning, Agents
Coding benchmarks
Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.
Writing & chat benchmarks
Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.
Hard reasoning benchmarks
Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.
Agents benchmarks
Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.