Skip to content
Model Pareto

Nex-N2.5-Mini

Nex AGIOpen weightshuggingface.co ↗
Input
—
Output
—
Speed
~120tok/sest.
Context
262Ktokens

Rankings

Self-hostingEstimated

How we estimate →

What it would cost to run these open weights yourself on rented GPUs. No API sells this model, so this is its price on the chart.

Hardware
1× H200 141GB
FP8 weights
Throughput
~8,390 tok/s
many requests batched
Price
$0.046–0.38 /M tok
busy → light use
On your own machine
RTX 5090 32GB (int4, 300+ tok/s single-stream)
single consumer GPU or Mac
Assumptions (5)
  • FP8 weights (35B params, 3B active per token) + 25% KV-cache headroom ≈ 44 GB
  • 1× H200 141GB at $3.59–$7.91/GPU-hour on-demand (2026-09-24)
  • ~8,390 output tok/s aggregate at batch 256 (bandwidth-bound); MoE compute scales with active params
  • Blended 3:1 input:output; prefill ~155,846 tok/s
  • Low = 75% utilization at the low GPU price; high = 20% at the high price

Benchmark scores

~ italic, dashed = no published score yet, estimated from related benchmarks.

Coding

Agents

Not comparable across labs (1)

Kept for reference, not counted in rankings: these use a lab-specific task set, answer key or scoring, so scores can't be compared fairly across labs. Only published scores are shown.