Compare models
Agents: quality vs price
Compared models are ringed; the other 170 ranked here are greyed.
Lower price is better. Pareto frontier: Gemma 4 E4B, Sarvam 30B, Qwen3.5 4B, Sarvam 105B, Ling 3.0 Flash, MiMo-V2.6-Flash, Qwen3.8 Flash-Next, GLM-5.3 Flash, Muse Spark 1.3, Claude Opus 5.5. 65 models are not plotted: Nex-N2.5-Pro, Atria Dawn Preview, Motif 3, K2 Horizon 375B A23B, Nex-N2.5-Mini, Solar Open2 250B, K2 Horizon MoVA 36B A4B, G9v3-39A5B, Apriel-v1.6-15B-Thinker, HyperCLOVA X SEED Think (32B), A.X-K2, EXAONE 4.5 33B, K2 Horizon 7B, Nex-N2-Pro, Tri-21B-Think, MiniCPM-V 4.6 1.3B, LongCat Flash Lite, MiniCPM5-1B, Qwen3.5 2B, Ling 3.0 Tiny, Nemotron Cascade 2 30B A3B, K-EXAONE 2.0 0803, K2 Horizon 3.7B, Solar Open 100B, Qwen3.5 0.8B, Command A+, MiniCPM5-2B, HyperNova 60B 2605, Nanbeige4.1-3B, K2-V2, Step3 VL 10B, INTELLECT-3, Falcon-H1R-7B, K2 Think V2, NVIDIA Nemotron 3 Nano 4B, Llama 3.1 Nemotron Ultra 253B v1, Gemma 4 E2B, North Mini Code, Devstral 2, LFM2.5-8B-A1B, Olmo 3.1 32B Instruct, LFM2.5-2.6B, Devstral Small 2, Hermes 4 - Llama-3.1 70B, LFM2.5-1.2B-Thinking, Exaone 4.0 1.2B, Granite 4.0 H 1B, Jamba Reasoning 3B, Granite 4.0 H 350M, Jamba 1.7 Large, Granite 4.0 350M, Jamba 1.7 Mini, LFM2 24B A2B, Granite 4.0 Micro, LFM2.5-1.2B-Instruct, EXAONE 4.0 32B, Gemma 3 270M, LFM2.5-VL-1.6B, K2 Horizon 0.9B, Olmo 3.1 32B Think, Olmo 3 7B Think, Molmo2-8B, Molmo 7B-D, Kimi Linear 48B A3B Instruct, Tiny Aya Global.
65 models with no price data — shown in the strip at the left edge
Nex-N2.5-Pro, Atria Dawn Preview, Motif 3, K2 Horizon 375B A23B, Nex-N2.5-Mini, Solar Open2 250B, K2 Horizon MoVA 36B A4B, G9v3-39A5B, Apriel-v1.6-15B-Thinker, HyperCLOVA X SEED Think (32B), A.X-K2, EXAONE 4.5 33B, K2 Horizon 7B, Nex-N2-Pro, Tri-21B-Think, MiniCPM-V 4.6 1.3B, LongCat Flash Lite, MiniCPM5-1B, Qwen3.5 2B, Ling 3.0 Tiny, Nemotron Cascade 2 30B A3B, K-EXAONE 2.0 0803, K2 Horizon 3.7B, Solar Open 100B, Qwen3.5 0.8B, Command A+, MiniCPM5-2B, HyperNova 60B 2605, Nanbeige4.1-3B, K2-V2, Step3 VL 10B, INTELLECT-3, Falcon-H1R-7B, K2 Think V2, NVIDIA Nemotron 3 Nano 4B, Llama 3.1 Nemotron Ultra 253B v1, Gemma 4 E2B, North Mini Code, Devstral 2, LFM2.5-8B-A1B, Olmo 3.1 32B Instruct, LFM2.5-2.6B, Devstral Small 2, Hermes 4 - Llama-3.1 70B, LFM2.5-1.2B-Thinking, Exaone 4.0 1.2B, Granite 4.0 H 1B, Jamba Reasoning 3B, Granite 4.0 H 350M, Jamba 1.7 Large, Granite 4.0 350M, Jamba 1.7 Mini, LFM2 24B A2B, Granite 4.0 Micro, LFM2.5-1.2B-Instruct, EXAONE 4.0 32B, Gemma 3 270M, LFM2.5-VL-1.6B, K2 Horizon 0.9B, Olmo 3.1 32B Think, Olmo 3 7B Think, Molmo2-8B, Molmo 7B-D, Kimi Linear 48B A3B Instruct, Tiny Aya Global
- Best-value frontier (nothing is both cheaper and better)
- Evidencestrong → weak
- Estimated from other categories
- Hover a dot for details · click a lab to highlight it
- K2 Horizon 7B
- Quality
- 40.5
- Rank
- #60/173
- Price
- —
- Speed
- —
vs Gemini 3.5 Flash-Lite: −0.1 quality · price n/a
- Gemini 3.5 Flash-Lite
- Quality
- 40.6
- Rank
- #59/173
- Price
- $0.85/M tok
- Speed
- 351tok/s
In$0.30Out$2.50/M tokvs K2 Horizon 7B: +0.1 quality · price n/a
- Nex-N2-Pro
- Quality
- 40.3
- Rank
- #61/173
- Price
- —
- Speed
- —
vs Gemini 3.5 Flash-Lite: −0.3 quality · price n/a
Overview
Quality by category
0–100 within each category (100 = best tracked model). Rank is among all models ranked there. Click a row to focus it.
Agents benchmarks
Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.
Benchmarks in other categories (4)· Coding, Writing & chat, Vision, Hard reasoning
Coding benchmarks
Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.
Writing & chat benchmarks
Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.
Vision benchmarks
Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.
Hard reasoning benchmarks
Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.