Qwen3.8 27B MXFP4
- Input
- —
- Output
- —
- Speed
- ~82tok/sest.
- Context
- 262Ktokens
Rankings
Self-hostingEstimated
How we estimate →What it would cost to run these open weights yourself on rented GPUs. No API sells this model, so this is its price on the chart.
- Hardware
- 1× B200 180GB
- MXFP4 weights
- Throughput
- ~5,953 tok/s
- many requests batched
- Price
- $0.14–1.22 /M tok
- busy → light use
- On your own machine
- RTX 5090 32GB (MXFP4, ~85 tok/s single-stream)
- single consumer GPU or Mac
Assumptions (5)
- MXFP4 weights (27.8B params) + 25% KV-cache headroom ≈ 18 GB
- 1× B200 180GB at $5.98–$14.24/GPU-hour on-demand (2026-09-24)
- ~5,953 output tok/s aggregate at batch 170 (bandwidth-bound)
- Blended 3:1 input:output; prefill ~38,242 tok/s
- Low = 75% utilization at the low GPU price; high = 20% at the high price
Benchmark scores
~ italic, dashed = no published score yet, estimated from related benchmarks.
Hard reasoning
GPQA DiamondLab-reported
#41 of 18387.7%
- #1GPT-6 Astra96.1%
- #2Gemini 3.8 Flash95.3%
- #40Qwen3.8 27B NVFP488.0%
- #41Qwen3.8 27B MXFP487.7%
- #42GPT-5.4 mini87.5%
huggingface.co · 2026-10-02
Humanity's Last ExamEstimated
~#54 of 185 (estimated)~26.4%
- #1Claude Opus 5.561.4%
- #2Claude Fable 5.159.1%
- #53MiMo-V2.527.2%
- ~#54Qwen3.8 27B MXFP4Estimated26.4%
- #54Solar Mini 425.8%
No published score — estimated from related benchmarks
AA Intelligence IndexEstimated
~#32 of 86 (estimated)~27.2
- #1Claude Opus 5.557.6
- #2Claude Sonnet 5.556.0
- #31MiniMax-M329.2
- ~#32Qwen3.8 27B MXFP4Estimated27.2
- #32Quasar 438B26.7
No published score — estimated from related benchmarks