Compare models
Writing & chat: quality vs price
Compared models are ringed; the other 143 ranked here are greyed.
Lower price is better. Pareto frontier: Gemma 4 E4B, Qwen3.5 4B, Gemma 4 12B, Qwen3.5 9B, Gemma 4 26B A4B, GPT-6 Luna, Qwen3.8 Flash-Next, MiMo-V2.6-Pro, Step 5 Preview, GLM-5.3, Kimi K3, Claude Opus 5.5, Claude Opus 5, Claude Fable 5.1. 52 models are not plotted: Intern-S2-Preview (35B-A3B), Xing4.0-29B-A4B, Nex-N2-Pro, Nemotron Cascade 2 30B A3B, Apriel-v1.6-15B-Thinker, Command A+, EXAONE 4.5 33B, HyperNova 60B 2605, DiffusionGemma 26B A4B, K2 Think V2, K2-V2, Solar Open 100B, Step3 VL 10B, Tri-21B-Think, Falcon-H1R-7B, North Mini Code, NVIDIA Nemotron 3 Nano 4B, LFM2.5-8B-A1B, LongCat Flash Lite, MiniCPM5-1B, Jamba Reasoning 3B, LFM2 24B A2B, Nanbeige4.1-3B, HyperCLOVA X SEED Think (32B), INTELLECT-3, Olmo 3 7B Think, LFM2.5-1.2B-Instruct, Llama 3.1 Nemotron Ultra 253B v1, LFM2.5-1.2B-Thinking, Devstral 2, Olmo 3.1 32B Instruct, Olmo 3.1 32B Think, Gemma 4 E2B, Qwen3.5 2B, EXAONE 4.0 32B, Devstral Small 2, Jamba 1.7 Large, MiniCPM-V 4.6 1.3B, LFM2.5-VL-1.6B, Hermes 4 - Llama-3.1 70B, Molmo2-8B, Jamba 1.7 Mini, Kimi Linear 48B A3B Instruct, Exaone 4.0 1.2B, Granite 4.0 H 1B, Granite 4.0 Micro, Qwen3.5 0.8B, Molmo 7B-D, Tiny Aya Global, Granite 4.0 H 350M, Granite 4.0 350M, Gemma 3 270M.
52 models with no price data — shown in the strip at the left edge
Intern-S2-Preview (35B-A3B), Xing4.0-29B-A4B, Nex-N2-Pro, Nemotron Cascade 2 30B A3B, Apriel-v1.6-15B-Thinker, Command A+, EXAONE 4.5 33B, HyperNova 60B 2605, DiffusionGemma 26B A4B, K2 Think V2, K2-V2, Solar Open 100B, Step3 VL 10B, Tri-21B-Think, Falcon-H1R-7B, North Mini Code, NVIDIA Nemotron 3 Nano 4B, LFM2.5-8B-A1B, LongCat Flash Lite, MiniCPM5-1B, Jamba Reasoning 3B, LFM2 24B A2B, Nanbeige4.1-3B, HyperCLOVA X SEED Think (32B), INTELLECT-3, Olmo 3 7B Think, LFM2.5-1.2B-Instruct, Llama 3.1 Nemotron Ultra 253B v1, LFM2.5-1.2B-Thinking, Devstral 2, Olmo 3.1 32B Instruct, Olmo 3.1 32B Think, Gemma 4 E2B, Qwen3.5 2B, EXAONE 4.0 32B, Devstral Small 2, Jamba 1.7 Large, MiniCPM-V 4.6 1.3B, LFM2.5-VL-1.6B, Hermes 4 - Llama-3.1 70B, Molmo2-8B, Jamba 1.7 Mini, Kimi Linear 48B A3B Instruct, Exaone 4.0 1.2B, Granite 4.0 H 1B, Granite 4.0 Micro, Qwen3.5 0.8B, Molmo 7B-D, Tiny Aya Global, Granite 4.0 H 350M, Granite 4.0 350M, Gemma 3 270M
- Best-value frontier (nothing is both cheaper and better)
- Evidencestrong → weak
- Estimated from other categories
- Hover a dot for details · click a lab to highlight it
- Gemma 4 31B
- Quality
- 63.5
- Rank
- #39/146
- Price
- $0.21/M tok
- Speed
- 35tok/s
In$0.14▾Out$0.40▾/M tokvs Xing4.0-29B-A4B: −3.9 quality · price n/a
- Xing4.0-29B-A4B
- Quality
- 67.4
- Rank
- #32/146
- Price
- —
- Speed
- —
vs Gemma 4 31B: +3.9 quality · price n/a
- Nemotron 3 Ultra
- Quality
- 60.7
- Rank
- #46/146
- Price
- $1.07/M tok
- Speed
- 215tok/s
In$0.60Out$2.50/M tokvs Xing4.0-29B-A4B: −6.7 quality · price n/a
Overview
Quality by category
0–100 within each category (100 = best tracked model). Rank is among all models ranked there. Click a row to focus it.
Writing & chat benchmarks
Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.
Benchmarks in other categories (4)· Coding, Vision, Hard reasoning, Agents
Coding benchmarks
Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.
Vision benchmarks
Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.
Hard reasoning benchmarks
Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.
Agents benchmarks
Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.