Kimi-K3-INT4
- Input
- —
- Output
- —
- Speed
- ~67tok/sest.
- Context
- 1Mtokens
Rankings
Self-hostingEstimated
How we estimate →What it would cost to run these open weights yourself on rented GPUs. No API sells this model, so this is its price on the chart.
- Hardware
- 16× B200 180GB
- INT4 weights
- Throughput
- ~21,178 tok/s
- many requests batched
- Price
- $0.61–5.44 /M tok
- busy → light use
- On your own machine
- Too large
- needs data-center GPUs
Assumptions (5)
- INT4 weights (2800B params, 104B active per token) + 40% KV-cache headroom ≈ 1960 GB
- 16× B200 180GB at $5.98–$14.24/GPU-hour on-demand (2026-09-24)
- ~21,178 output tok/s aggregate at batch 713 (bandwidth-bound); MoE compute scales with active params
- Blended 3:1 input:output; prefill ~139,024 tok/s
- Low = 75% utilization at the low GPU price; high = 20% at the high price
Benchmark scores
~ italic, dashed = no published score yet, estimated from related benchmarks.
Hard reasoning
GPQA DiamondLab-reported
#20 of 18392.0%
- #1GPT-6 Astra96.1%
- #2Gemini 3.8 Flash95.3%
- #19GLM-5.3 Flash NVFP492.1%
- #20Kimi-K3-INT492.0%
- #21GLM-5.391.7%
huggingface.co · 2026-10-02
Humanity's Last ExamEstimated
~#46 of 185 (estimated)~29.7%
- #1Claude Opus 5.561.4%
- #2Claude Fable 5.159.1%
- #45Inkling31.9%
- ~#46Kimi-K3-INT4Estimated29.7%
- #46A.X-K229.6%
No published score — estimated from related benchmarks
AIME (latest)Estimated
~#2 of 32 (estimated)~98.9%
No published score — estimated from related benchmarks
AA Intelligence IndexEstimated
~#29 of 86 (estimated)~32.3
- #1Claude Opus 5.557.6
- #2Claude Sonnet 5.556.0
- #28Qwen3.8 27B33.7
- ~#29Kimi-K3-INT4Estimated32.3
- #29K2 Horizon 375B A23B30.5
No published score — estimated from related benchmarks