Skip to content
Model Pareto

Compare models

Side by side: quality per category, price, speed, and every benchmark cell with where it came from. Pick up to four.
Ternary Bonsai 27BNex-N2-ProK2 Horizon 375B A23B
Focus

Coding: quality vs price

Compared models are ringed; the other 111 ranked here are greyed.

Lower price is better. Pareto frontier: Granite 4.2 3B, NVIDIA Nemotron 3 Nano 30B A3B, gpt-oss-20b, Ling-3.0-flash-VL, MiMo-V2.6-Flash, Qwen3.8-Omni-Flash, GLM-5.3 Flash, DeepSeek V4.1 Flash, MiMo-V2.6-Pro, Claude Sonnet 5.5, Claude Opus 5.5. 36 models are not plotted: GLM-5.3 Flash NVFP4, Ternary Bonsai 2 27B, DeepSeek V4.1 Flash NVFP4, GLM-5.3 NVFP4, Nex-N2.5-Pro, Xing4.0-29B-A4B, Atria Dawn Preview, Ternary Bonsai 27B, Maple-Preview, Bonsai 27B, Qwen3.8 27B NVFP4, Intern-S2-397B, Solar Open2 250B, Phi-4-reasoning-plus, Motif 3, Nex-N2-Pro, Nex-N2.5-Mini, A.X-K2, K2 Horizon 375B A23B, K-EXAONE 2.0 0803, K2 Horizon MoVA 36B A4B, Intern-S2-Preview (35B-A3B), Command A+, North Mini Code, G9v3-39A5B, Falcon-H1-34B-Instruct, Qwen3.8 Flash-Next NVFP4, Devstral 2, K2 Horizon 7B, K2-Horizon-32B, Devstral Small 2, MiniCPM5-2B, Ling 3.0 Tiny, K2 Horizon 3.7B, K2 Horizon 0.9B, LFM2.5-2.6B.

36 models with no price data — shown in the strip at the left edge

GLM-5.3 Flash NVFP4, Ternary Bonsai 2 27B, DeepSeek V4.1 Flash NVFP4, GLM-5.3 NVFP4, Nex-N2.5-Pro, Xing4.0-29B-A4B, Atria Dawn Preview, Ternary Bonsai 27B, Maple-Preview, Bonsai 27B, Qwen3.8 27B NVFP4, Intern-S2-397B, Solar Open2 250B, Phi-4-reasoning-plus, Motif 3, Nex-N2-Pro, Nex-N2.5-Mini, A.X-K2, K2 Horizon 375B A23B, K-EXAONE 2.0 0803, K2 Horizon MoVA 36B A4B, Intern-S2-Preview (35B-A3B), Command A+, North Mini Code, G9v3-39A5B, Falcon-H1-34B-Instruct, Qwen3.8 Flash-Next NVFP4, Devstral 2, K2 Horizon 7B, K2-Horizon-32B, Devstral Small 2, MiniCPM5-2B, Ling 3.0 Tiny, K2 Horizon 3.7B, K2 Horizon 0.9B, LFM2.5-2.6B

  • Best-value frontier (nothing is both cheaper and better)
  • Evidencestrong → weak
  • Estimated from other categories
  • Ternary Bonsai 27B
    Quality
    59.8
    Rank
    #28/114
    Price
    —
    Speed
    —

    vs Nex-N2-Pro: +20.4 quality · price n/a

  • Nex-N2-Pro
    Quality
    39.4
    Rank
    #51/114
    Price
    —
    Speed
    —

    vs Ternary Bonsai 27B: −20.3 quality · price n/a

  • K2 Horizon 375B A23B
    Quality
    35.3
    Rank
    #64/114
    Price
    —
    Speed
    —

    vs Ternary Bonsai 27B: −24.5 quality · price n/a

Overview

Lab
PrismML
Nex AGI
MBZUAI Institute of Foundation Models
Released
Jul 4, 2026
Jun 2, 2026
Sep 3, 2026
Weights
Open
Open
Open
Context
tokens
262K
262K
524KBest

Quality by category

0–100 within each category (100 = best tracked model). Rank is among all models ranked there. Click a row to focus it.

59.8±10.3Best
#28/114Lab-reported
39.4±10.0
#51/114Verified
35.3±8.0
#64/114Verified
63.4±9.7Best
#47/155Lab-reported
60.8±9.5
#52/155Verified
Not ranked
56.7±8.7
#41/90Lab-reported
Not ranked
Not ranked
65.8±9.9
#42/193Lab-reported
68.1±5.9Best
#41/193Verified
65.7±3.7
#43/193Verified
Not ranked
40.3±4.2
#64/176Verified
56.3±5.8Best
#32/176Verified

Coding benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

~#9/15~52.9Estimated
estimated from related benchmarks
~last/15below measured range (<41.1)Estimated
estimated from related benchmarks
~last/15below measured range (<41.1)Estimated
estimated from related benchmarks
Benchmarks in other categories (4)

Writing & chat benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

Vision benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

Hard reasoning benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

Agents benchmarks

Rank among models with a published score, and the leaderboard around each model. ~ italic = no published score, estimated from related benchmarks.

—
~last/5below measured range (<47.5%)Estimated
estimated from related benchmarks
~last/5below measured range (<47.5%)Estimated
estimated from related benchmarks