Compare models
Vision: quality vs price
Compared models are ringed; the other 87 ranked here are greyed.
Lower price is better. Pareto frontier: Gemma 4 E4B, Qwen3.5 4B, Ling-3.0-flash-VL, MiMo-V2.5, GPT-6 Luna, Qwen3.8-Omni-Flash, Gemini 3.8 Flash, GPT-6 Sol, Claude Opus 5.5, GPT-6 Astra. 23 models are not plotted: Intern-S2-397B, Qwen3.8 Flash-Next NVFP4, Ternary Bonsai 2 27B, GLM-5.3 Flash NVFP4, Qwen3.8 27B NVFP4, DeepSeek V4.1 Flash NVFP4, Intern-S2-Preview (35B-A3B), Ternary Bonsai 27B, EXAONE 4.5 33B, Bonsai 27B, DiffusionGemma 26B A4B, Step3 VL 10B, Command A+, Devstral Small 2, Qwen3.5 2B, Gemma 4 E2B, Llama 3.2 Instruct 90B (Vision), MiniCPM-V 4.6 1.3B, Molmo2-8B, LFM2.5-VL-1.6B, Qwen3.5 0.8B, Molmo 7B-D, Phi-4 Multimodal Instruct.
23 models with no price data — shown in the strip at the left edge
Intern-S2-397B, Qwen3.8 Flash-Next NVFP4, Ternary Bonsai 2 27B, GLM-5.3 Flash NVFP4, Qwen3.8 27B NVFP4, DeepSeek V4.1 Flash NVFP4, Intern-S2-Preview (35B-A3B), Ternary Bonsai 27B, EXAONE 4.5 33B, Bonsai 27B, DiffusionGemma 26B A4B, Step3 VL 10B, Command A+, Devstral Small 2, Qwen3.5 2B, Gemma 4 E2B, Llama 3.2 Instruct 90B (Vision), MiniCPM-V 4.6 1.3B, Molmo2-8B, LFM2.5-VL-1.6B, Qwen3.5 0.8B, Molmo 7B-D, Phi-4 Multimodal Instruct
- Best-value frontier (nothing is both cheaper and better)
- Evidencestrong → weak
- Estimated from other categories
- Hover a dot for details · click a lab to highlight it
- GLM-5.3 Flash NVFP4
- Quality
- 68.8
- Rank
- #22/90
- Price
- —
- Speed
- —
vs Ternary Bonsai 2 27B: −0.2 quality · price n/a
- Step 5 Preview
- Quality
- 66.1
- Rank
- #26/90
- Price
- $1.43/M tok
- Speed
- 85tok/s
In$1.00Out$2.70/M tokvs Ternary Bonsai 2 27B: −2.9 quality · price n/a
- Ternary Bonsai 2 27B
- Quality
- 69.0
- Rank
- #20/90
- Price
- —
- Speed
- —
vs GLM-5.3 Flash NVFP4: +0.2 quality · price n/a
Overview
Quality by category
0–100 within each category (100 = best tracked model). Rank is among all models ranked there. Click a row to focus it.
Vision benchmarks
Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.
Benchmarks in other categories (4)· Coding, Writing & chat, Hard reasoning, Agents
Coding benchmarks
Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.
Writing & chat benchmarks
Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.
Hard reasoning benchmarks
Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.
Agents benchmarks
Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.