Skip to content
Model Pareto

GLM-5.3 NVFP4

NVIDIAOpen weightsNVFP4quantized from GLM-5.3huggingface.co ↗
Input
—
Output
—
Speed
~140tok/sest.
Context
1Mtokens

Rankings

Self-hostingEstimated

How we estimate →

What it would cost to run these open weights yourself on rented GPUs. No API sells this model, so this is its price on the chart.

Hardware
8× H200 141GB
NVFP4 weights
Throughput
~16,000 tok/s
many requests batched
Price
$0.25–2.08 /M tok
busy → light use
On your own machine
Too large
needs data-center GPUs
Assumptions (5)
  • NVFP4 weights (744B params, 40B active per token) + 40% KV-cache headroom ≈ 586 GB
  • 8× H200 141GB at $3.59–$7.91/GPU-hour on-demand (2026-09-24)
  • ~16,000 output tok/s aggregate at batch 575 (bandwidth-bound); MoE compute scales with active params
  • Blended 3:1 input:output; prefill ~93,508 tok/s
  • Low = 75% utilization at the low GPU price; high = 20% at the high price

Benchmark scores

~ italic, dashed = no published score yet, estimated from related benchmarks.

Coding

Writing & chat

Hard reasoning