Skip to content
Model Pareto

GLM-5.3-MXFP4

Red Hat AIOpen weightsMXFP4quantized from GLM-5.3huggingface.co ↗
Input
—
Output
—
Speed
~78tok/sest.
Context
1Mtokens

Red Hat AI MXFP4 (W4A4, group 32, E8M0 scales) quant of GLM-5.3; ~465 GB on disk. bits 4.25 = 4-bit e2m1 + 8-bit shared scale per 32 weights, matching the other MXFP4 entries. Context from config.json max_position_embeddings 1048576 (card repo and base).

Rankings

Self-hostingEstimated

How we estimate →

What it would cost to run these open weights yourself on rented GPUs. No API sells this model, so this is its price on the chart.

Hardware
8× H200 141GB
MXFP4 weights
Throughput
~16,630 tok/s
many requests batched
Price
$0.25–2.03 /M tok
busy → light use
On your own machine
Mac Studio M5 Ultra 512GB (MXFP4, ~40 tok/s single-stream)
single consumer GPU or Mac
Assumptions (5)
  • MXFP4 weights (744B params, 40B active per token) + 40% KV-cache headroom ≈ 553 GB
  • 8× H200 141GB at $3.59–$7.91/GPU-hour on-demand (2026-09-24)
  • ~16,630 output tok/s aggregate at batch 598 (bandwidth-bound); MoE compute scales with active params
  • Blended 3:1 input:output; prefill ~93,508 tok/s
  • Low = 75% utilization at the low GPU price; high = 20% at the high price

Benchmark scores

~ italic, dashed = no published score yet, estimated from related benchmarks.