Skip to content
Model Pareto

Compare models

Side by side: quality per category, price, speed, and every benchmark cell with where it came from. Pick up to four.
K2 Horizon 375B A23BNex-N2-ProTernary Bonsai 27B
Focus

Coding: quality vs price

Compared models are ringed; the other 111 ranked here are greyed.

Lower price is better. Pareto frontier: Granite 4.2 3B, NVIDIA Nemotron 3 Nano 30B A3B, gpt-oss-20b, Ling-3.0-flash-VL, MiMo-V2.6-Flash, Qwen3.8-Omni-Flash, GLM-5.3 Flash, DeepSeek V4.1 Flash, MiMo-V2.6-Pro, Claude Sonnet 5.5, Claude Opus 5.5. 36 models are not plotted: GLM-5.3 Flash NVFP4, Ternary Bonsai 2 27B, DeepSeek V4.1 Flash NVFP4, GLM-5.3 NVFP4, Nex-N2.5-Pro, Xing4.0-29B-A4B, Atria Dawn Preview, Ternary Bonsai 27B, Maple-Preview, Bonsai 27B, Qwen3.8 27B NVFP4, Intern-S2-397B, Solar Open2 250B, Phi-4-reasoning-plus, Motif 3, Nex-N2-Pro, Nex-N2.5-Mini, A.X-K2, K2 Horizon 375B A23B, K-EXAONE 2.0 0803, K2 Horizon MoVA 36B A4B, Intern-S2-Preview (35B-A3B), Command A+, North Mini Code, G9v3-39A5B, Falcon-H1-34B-Instruct, Qwen3.8 Flash-Next NVFP4, Devstral 2, K2 Horizon 7B, K2-Horizon-32B, Devstral Small 2, MiniCPM5-2B, Ling 3.0 Tiny, K2 Horizon 3.7B, K2 Horizon 0.9B, LFM2.5-2.6B.

36 models with no price data — shown in the strip at the left edge

GLM-5.3 Flash NVFP4, Ternary Bonsai 2 27B, DeepSeek V4.1 Flash NVFP4, GLM-5.3 NVFP4, Nex-N2.5-Pro, Xing4.0-29B-A4B, Atria Dawn Preview, Ternary Bonsai 27B, Maple-Preview, Bonsai 27B, Qwen3.8 27B NVFP4, Intern-S2-397B, Solar Open2 250B, Phi-4-reasoning-plus, Motif 3, Nex-N2-Pro, Nex-N2.5-Mini, A.X-K2, K2 Horizon 375B A23B, K-EXAONE 2.0 0803, K2 Horizon MoVA 36B A4B, Intern-S2-Preview (35B-A3B), Command A+, North Mini Code, G9v3-39A5B, Falcon-H1-34B-Instruct, Qwen3.8 Flash-Next NVFP4, Devstral 2, K2 Horizon 7B, K2-Horizon-32B, Devstral Small 2, MiniCPM5-2B, Ling 3.0 Tiny, K2 Horizon 3.7B, K2 Horizon 0.9B, LFM2.5-2.6B

  • Best-value frontier (nothing is both cheaper and better)
  • Evidencestrong → weak
  • Estimated from other categories
  • K2 Horizon 375B A23B
    Quality
    35.3
    Rank
    #64/114
    Price
    —
    Speed
    —

    vs Ternary Bonsai 27B: −24.5 quality · price n/a

  • Nex-N2-Pro
    Quality
    39.4
    Rank
    #51/114
    Price
    —
    Speed
    —

    vs Ternary Bonsai 27B: −20.3 quality · price n/a

  • Ternary Bonsai 27B
    Quality
    59.8
    Rank
    #28/114
    Price
    —
    Speed
    —

    vs Nex-N2-Pro: +20.4 quality · price n/a

Overview

Lab
MBZUAI Institute of Foundation Models
Nex AGI
PrismML
Released
Sep 3, 2026
Jun 2, 2026
Jul 4, 2026
Weights
Open
Open
Open
Context
tokens
524KBest
262K
262K

Quality by category

0–100 within each category (100 = best tracked model). Rank is among all models ranked there. Click a row to focus it.

35.3±8.0
#64/114Verified
39.4±10.0
#51/114Verified
59.8±10.3Best
#28/114Lab-reported
Not ranked
60.8±9.5
#52/155Verified
63.4±9.7Best
#47/155Lab-reported
Not ranked
Not ranked
56.7±8.7
#41/90Lab-reported
65.7±3.7
#43/193Verified
68.1±5.9Best
#41/193Verified
65.8±9.9
#42/193Lab-reported
56.3±5.8Best
#32/176Verified
40.3±4.2
#64/176Verified
Not ranked

Coding benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

~last/15below measured range (<41.1)Estimated
estimated from related benchmarks
~last/15below measured range (<41.1)Estimated
estimated from related benchmarks
~#9/15~52.9Estimated
estimated from related benchmarks
Benchmarks in other categories (4)

Writing & chat benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

Vision benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

Hard reasoning benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

Agents benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

~last/5below measured range (<47.5%)Estimated
estimated from related benchmarks
~last/5below measured range (<47.5%)Estimated
estimated from related benchmarks
—