Hy3
- Input
- $0.14/M tok
- Output
- $0.58/M tok
- Blended
- $0.25/M tok · 3:1
- Speed
- 86tok/s
- Context
- 256Ktokens
Open weights.
Rankings
Coding
Mixed#61 / 103rank by quality34.4±4.3quality / 1005/7 benchmarks measuredWriting & chat
Verified#30 / 146rank by quality67.7±6.0quality / 1002/5 benchmarks measuredHard reasoning
Verified#41 / 182rank by quality64.1±3.7quality / 1003/4 benchmarks measuredAgents
Mixed#48 / 173rank by quality45.1±4.6quality / 1004/7 benchmarks measured
Self-hostingEstimated
How we estimate →What it would cost to run these open weights yourself on rented GPUs, compared with the API price above.
- Hardware
- 4× B200 180GB
- FP8 weights
- Throughput
- ~16,423 tok/s
- many requests batched
- Price
- $0.17–1.50 /M tok
- busy → light use
- On your own machine
- Mac Studio M5 Ultra 512GB (int4, ~80 tok/s single-stream)
- single consumer GPU or Mac
Assumptions (5)
- FP8 weights (295B params, 21B active per token) + 25% KV-cache headroom ≈ 369 GB
- 4× B200 180GB at $5.98–$14.24/GPU-hour on-demand (2026-09-24)
- ~16,423 output tok/s aggregate at batch 470 (bandwidth-bound); MoE compute scales with active params
- Blended 3:1 input:output; prefill ~202,500 tok/s
- Low = 75% utilization at the low GPU price; high = 20% at the high price
Benchmark scores
~ italic, dashed = no published score yet, estimated from related benchmarks.
Coding
LiveCodeBenchEstimated
~#7 of 20 (estimated)~78.0%
- #1Qwen3.8-Omni-Flash92.6%
- #2Qwen3.8 Flash-Next91.9%
- #6Gemma 4 31B80.0%
- ~#7Hy3Estimated78.0%
- #7gpt-oss-20b77.7%
No published score — estimated from related benchmarks
AA Coding Agent IndexEstimated
~last of 15 (estimated)below measured range (<41.1)
No published score — estimated from related benchmarks
Writing & chat
IFBenchEstimated
~#33 of 126 (estimated)~67.2%
- #1MiniMax-M382.9%
- #2Qwen3.8 2.4T A95B82.8%
- #32Step 3.7 Flash67.3%
- ~#33Hy3Estimated67.2%
- #33MiMo-V2.567.1%
No published score — estimated from related benchmarks
LMArena Text, Non-English (Elo)Verified
#22 of 531440 Elo
- #1Claude Opus 5.51501
- #2Claude Fable 5.1Few votes1496
- #21Qwen3.7 Plus1443
- #22Hy31440
- #23GPT-5.4 mini1439
lmarena.ai · 2026-09-24
LMArena Text (Elo)Verified
#19 of 531456 Elo
- #1Claude Opus 5.51509
- #2Claude Fable 5.11501
- #19Gemini 3.5 Flash-Lite1456
- #19Hy31456
- #19Qwen3.7 Plus1456
lmarena.ai · 2026-09-24
EQ-Bench Creative Writing v3 (Elo)Estimated
~#13 of 26 (estimated)~1832 Elo
- #1GPT-6 Astra2173
- #2Claude Fable 5.12162
- #12Qwen3.8 2.4T A95B1843
- ~#13Hy3Estimated1832
- #13Muse Glimmer1798
No published score — estimated from related benchmarks
Hard reasoning
AIME (latest)Estimated
~#5 of 25 (estimated)~92.8%
No published score — estimated from related benchmarks
Agents
τ²-bench (Telecom)Estimated
~#36 of 116 (estimated)~77.2%
- #1Step 3.7 Flash98.5%
- #2Gemini 3.1 Pro95.6%
- #35EXAONE 4.5 33B78.1%
- ~#36Hy3Estimated77.2%
- #36GPT-5.4 nano76.0%
No published score — estimated from related benchmarks
OSWorld-VerifiedEstimated
~#10 of 12 (estimated)~70.1%
No published score — estimated from related benchmarks
OSWorld 2.0Estimated
~last of 5 (estimated)below measured range (<47.5%)
No published score — estimated from related benchmarks
GDPval-AA (Elo)Verified
#34 of 861045 Elo
- #1Claude Opus 5.51846
- #2Claude Fable 5.11735
- #33Grok Build 0.1 06161052
- #34Hy31045
- #35Kimi K2.7 Code1025
artificialanalysis.ai · 2026-09-24