Skip to content
Model Pareto

Compare models

Side by side: quality per category, price, speed, and every benchmark cell with where it came from. Pick up to four.
GLM-5.3-MXFP4Nex-N2-ProTernary Bonsai 27B
Focus

Hard reasoning: quality vs price

Compared models are ringed; the other 197 ranked here are greyed.

Lower price is better. Pareto frontier: Gemma 4 E4B, Sarvam 30B, Qwen3.5 4B, Sarvam 105B, Ling-3.0-flash-VL, MiMo-V2.5, Ternary Bonsai 2 27B, Claude Haiku 5.5, MiMo-V2.6-Pro, Muse Spark 1.3, Claude Sonnet 5.5, Claude Opus 5.5. 81 models are not plotted: DeepSeek V4.1 Flash NVFP4, GLM-5.3 Flash NVFP4, GLM-5.3 NVFP4, Qwen3.8 Flash-Next NVFP4, DeepSeek-V4-Pro-0813-nvfp4-DSpark, Kimi-K3-INT4, Qwen3.8 27B NVFP4, Motif 3, Nex-N2-Pro, GLM-5.3-MXFP4, Ternary Bonsai 27B, K2 Horizon 375B A23B, Qwen3.8 27B MXFP4, Solar Open2 250B, Bonsai 27B, Xing4.0-29B-A4B, A.X-K2, K2 Horizon MoVA 36B A4B, K2-Horizon-32B, Maple-Preview, Intern-S2-Preview (35B-A3B), K-EXAONE 2.0 0803, G9v3-39A5B, K2 Horizon 7B, Phi-4-reasoning-plus, EXAONE 4.5 33B, Nanbeige4.1-3B, HyperNova 60B 2605, Nemotron Cascade 2 30B A3B, INTELLECT-3, Apriel-v1.6-15B-Thinker, K2 Horizon 3.7B, K2 Think V2, Step3 VL 10B, DiffusionGemma 26B A4B, North Mini Code, K2-V2, Falcon-H1R-7B, Ling 3.0 Tiny, Solar Open 100B, Llama 3.1 Nemotron Ultra 253B v1, MiniCPM5-2B, LongCat Flash Lite, HyperCLOVA X SEED Think (32B), Tri-21B-Think, EXAONE 4.0 32B, Olmo 3.1 32B Think, LFM2.5-2.6B, Devstral 2, LFM2.5-8B-A1B, Olmo 3.1 32B Instruct, Olmo 3 7B Think, Devstral Small 2, NVIDIA Nemotron 3 Nano 4B, Phi-4-mini-flash-reasoning, Qwen3.5 2B, Falcon-H1-34B-Instruct, Hermes 4 - Llama-3.1 70B, LFM2 24B A2B, Exaone 4.0 1.2B, Llama 3.2 Instruct 90B (Vision), Gemma 4 E2B, Molmo2-8B, LFM2.5-1.2B-Thinking, LFM2.5-1.2B-Instruct, MiniCPM5-1B, Jamba 1.7 Large, MiniCPM-V 4.6 1.3B, Kimi Linear 48B A3B Instruct, Jamba Reasoning 3B, Jamba 1.7 Mini, Phi-4 Multimodal Instruct, Tiny Aya Global, LFM2.5-VL-1.6B, Granite 4.0 H 350M, K2 Horizon 0.9B, Qwen3.5 0.8B, Granite 4.0 350M, Granite 4.0 H 1B, Molmo 7B-D, Gemma 3 270M.

81 models with no price data — shown in the strip at the left edge

DeepSeek V4.1 Flash NVFP4, GLM-5.3 Flash NVFP4, GLM-5.3 NVFP4, Qwen3.8 Flash-Next NVFP4, DeepSeek-V4-Pro-0813-nvfp4-DSpark, Kimi-K3-INT4, Qwen3.8 27B NVFP4, Motif 3, Nex-N2-Pro, GLM-5.3-MXFP4, Ternary Bonsai 27B, K2 Horizon 375B A23B, Qwen3.8 27B MXFP4, Solar Open2 250B, Bonsai 27B, Xing4.0-29B-A4B, A.X-K2, K2 Horizon MoVA 36B A4B, K2-Horizon-32B, Maple-Preview, Intern-S2-Preview (35B-A3B), K-EXAONE 2.0 0803, G9v3-39A5B, K2 Horizon 7B, Phi-4-reasoning-plus, EXAONE 4.5 33B, Nanbeige4.1-3B, HyperNova 60B 2605, Nemotron Cascade 2 30B A3B, INTELLECT-3, Apriel-v1.6-15B-Thinker, K2 Horizon 3.7B, K2 Think V2, Step3 VL 10B, DiffusionGemma 26B A4B, North Mini Code, K2-V2, Falcon-H1R-7B, Ling 3.0 Tiny, Solar Open 100B, Llama 3.1 Nemotron Ultra 253B v1, MiniCPM5-2B, LongCat Flash Lite, HyperCLOVA X SEED Think (32B), Tri-21B-Think, EXAONE 4.0 32B, Olmo 3.1 32B Think, LFM2.5-2.6B, Devstral 2, LFM2.5-8B-A1B, Olmo 3.1 32B Instruct, Olmo 3 7B Think, Devstral Small 2, NVIDIA Nemotron 3 Nano 4B, Phi-4-mini-flash-reasoning, Qwen3.5 2B, Falcon-H1-34B-Instruct, Hermes 4 - Llama-3.1 70B, LFM2 24B A2B, Exaone 4.0 1.2B, Llama 3.2 Instruct 90B (Vision), Gemma 4 E2B, Molmo2-8B, LFM2.5-1.2B-Thinking, LFM2.5-1.2B-Instruct, MiniCPM5-1B, Jamba 1.7 Large, MiniCPM-V 4.6 1.3B, Kimi Linear 48B A3B Instruct, Jamba Reasoning 3B, Jamba 1.7 Mini, Phi-4 Multimodal Instruct, Tiny Aya Global, LFM2.5-VL-1.6B, Granite 4.0 H 350M, K2 Horizon 0.9B, Qwen3.5 0.8B, Granite 4.0 350M, Granite 4.0 H 1B, Molmo 7B-D, Gemma 3 270M

  • Best-value frontier (nothing is both cheaper and better)
  • Evidencestrong → weak
  • Estimated from other categories
  • GLM-5.3-MXFP4
    Quality
    66.1
    Rank
    #45/200
    Price
    —
    Speed
    —

    vs Nex-N2-Pro: −1.8 quality · price n/a

  • Nex-N2-Pro
    Quality
    67.9
    Rank
    #44/200
    Price
    —
    Speed
    —

    vs GLM-5.3-MXFP4: +1.8 quality · price n/a

  • Ternary Bonsai 27B
    Quality
    65.8
    Rank
    #46/200
    Price
    —
    Speed
    —

    vs Nex-N2-Pro: −2.1 quality · price n/a

Overview

Lab
Red Hat AI
Nex AGI
PrismML
Released
Sep 15, 2026
Jun 2, 2026
Jul 4, 2026
Weights
Open
Open
Open
Context
tokens
1MBest
262K
262K

Quality by category

0–100 within each category (100 = best tracked model). Rank is among all models ranked there. Click a row to focus it.

Not ranked
38.4±10.0
#55/119Verified
60.3±10.4Best
#30/119Lab-reported
Not ranked
58.9±9.5
#53/156Verified
61.4±9.7Best
#47/156Lab-reported
Not ranked
Not ranked
56.3±8.7
#41/91Lab-reported
66.1±8.3
#45/200Lab-reported
67.9±5.9Best
#44/200Verified
65.8±9.9
#46/200Lab-reported
Not ranked
40.3±4.2
#65/178Verified
Not ranked

Hard reasoning benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

Benchmarks in other categories (4)

Coding benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

—
~last/18below measured range (<41.1)Estimated
estimated from related benchmarks
~#11/18~53.4Estimated
estimated from related benchmarks

Writing & chat benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

Vision benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

Agents benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.