GLM-5.3 NVFP4
- Input
- —
- Output
- —
- Speed
- ~140tok/sest.
- Context
- 1Mtokens
Rankings
Self-hostingEstimated
How we estimate →What it would cost to run these open weights yourself on rented GPUs. No API sells this model, so this is its price on the chart.
- Hardware
- 8× H200 141GB
- NVFP4 weights
- Throughput
- ~16,000 tok/s
- many requests batched
- Price
- $0.25–2.08 /M tok
- busy → light use
- On your own machine
- Too large
- needs data-center GPUs
Assumptions (5)
- NVFP4 weights (744B params, 40B active per token) + 40% KV-cache headroom ≈ 586 GB
- 8× H200 141GB at $3.59–$7.91/GPU-hour on-demand (2026-09-24)
- ~16,000 output tok/s aggregate at batch 575 (bandwidth-bound); MoE compute scales with active params
- Blended 3:1 input:output; prefill ~93,508 tok/s
- Low = 75% utilization at the low GPU price; high = 20% at the high price
Benchmark scores
~ italic, dashed = no published score yet, estimated from related benchmarks.
Coding
SWE-bench VerifiedEstimated
~#2 of 14 (estimated)~88.0%
No published score — estimated from related benchmarks
SWE-bench Pro (Public)Estimated
~#5 of 19 (estimated)~71.4%
- #1Claude Opus 5.589.9%
- #2Claude Sonnet 5.581.3%
- #4Claude Opus 579.2%
- ~#5GLM-5.3 NVFP4Estimated71.4%
- #5Intern-S2-397B68.5%
No published score — estimated from related benchmarks
DeepSWE v1.1Estimated
~#20 of 32 (estimated)~63.4%
- #1GPT-6.1 Sol75.2%
- #2DeepSeek V4.1 Flash74.2%
- #19GPT-6 Luna63.7%
- ~#20GLM-5.3 NVFP4Estimated63.4%
- #20GLM-5.3 Flash63.4%
No published score — estimated from related benchmarks
Terminal-Bench 4.0Estimated
~#15 of 85 (estimated)~27.3%
- #1Claude Sonnet 5.563.6%
- #2Claude Opus 5.559.6%
- #14GLM-5.3 Flash32.8%
- ~#15GLM-5.3 NVFP4Estimated27.3%
- #15DeepSeek V4.1 Flash26.8%
No published score — estimated from related benchmarks
LiveCodeBenchLab-reported
#14 of 2453.5%
- #1Qwen3.8-Omni-Flash92.6%
- #2Qwen3.8 Flash-Next91.9%
- #13Claude Haiku 4.561.5%
- #14GLM-5.3 NVFP453.5%
- #15Phi-4-reasoning-plus53.1%
huggingface.co · 2026-09-29
SciCodeLab-reported
#10 of 10258.0%
- #1Claude Opus 5.566.9%
- #2Claude Fable 5.163.1%
- #9Gemini 3.1 Pro58.7%
- #10GLM-5.3 NVFP458.0%
- #11GLM-5.3 Flash NVFP457.7%
huggingface.co · 2026-09-29
AA Coding Agent IndexEstimated
~#6 of 15 (estimated)~56.5
- #1Claude Opus 5.566.0
- #2Claude Fable 5.162.2
- #5GPT-6 Sol56.7
- ~#6GLM-5.3 NVFP4Estimated56.5
- #6Grok 4.756.3
No published score — estimated from related benchmarks
Writing & chat
IFBenchLab-reported
#43 of 13366.0%
- #1MiniMax-M382.9%
- #2Qwen3.8 2.4T A95B82.8%
- #41Nova 2.0 Omni66.2%
- #43GLM-5.3 NVFP466.0%
- #43Olmo 3.1 32B Think66.0%
huggingface.co · 2026-09-29
LMArena Text, Non-English (Elo)Estimated
~#20 of 54 (estimated)~1446 Elo
- #1Claude Opus 5.51501
- #2Claude Fable 5.1Few votes1496
- #18Grok 4.61447
- ~#20GLM-5.3 NVFP4Estimated1446
- #20Gemini 3.5 Flash-Lite1445
No published score — estimated from related benchmarks
LMArena Text (Elo)Estimated
~#18 of 54 (estimated)~1458 Elo
- #1Claude Opus 5.51509
- #2Claude Fable 5.11501
- #17Claude Sonnet 51461
- ~#18GLM-5.3 NVFP4Estimated1458
- #18GPT-6 Sol1457
No published score — estimated from related benchmarks
EQ-Bench Creative Writing v3 (Elo)Estimated
~#9 of 26 (estimated)~1988 Elo
- #1GPT-6 Astra2173
- #2Claude Fable 5.12162
- #8Grok 4.72007
- ~#9GLM-5.3 NVFP4Estimated1988
- #9Muse Spark 1.31906
No published score — estimated from related benchmarks
Hard reasoning
GPQA DiamondLab-reported
#15 of 17992.7%
- #1GPT-6 Astra96.1%
- #2Gemini 3.8 Flash95.3%
- #13Qwen3.8 Max (0902)92.8%
- #15GLM-5.3 NVFP492.7%
- #16GPT-5.6 Terra92.5%
huggingface.co · 2026-09-29
Humanity's Last ExamEstimated
~#44 of 182 (estimated)~29.7%
- #1Claude Opus 5.561.4%
- #2Claude Fable 5.159.1%
- #43Inkling31.9%
- ~#44GLM-5.3 NVFP4Estimated29.7%
- #44A.X-K229.6%
No published score — estimated from related benchmarks
AA Intelligence IndexEstimated
~#27 of 78 (estimated)~33.1
- #1Claude Opus 5.557.6
- #2Claude Sonnet 5.556.0
- #26Qwen3.8 27B33.7
- ~#27GLM-5.3 NVFP4Estimated33.1
- #27K2 Horizon 375B A23B30.5
No published score — estimated from related benchmarks