Atria Dawn Preview
- Input
- —
- Output
- —
- Speed
- ~81tok/sest.
- Context
- 262Ktokens
Agentic model from Shanghai AI Lab built on GLM-5.2 (744B MoE). No API price found. Card also reports self-run Terminal-Bench 2.1, AutomationBench, GDPval and tau3-Banking, not mapped (our ids hold AA / Terminal-Bench 4.0 runs).
Rankings
Self-hostingEstimated
How we estimate →What it would cost to run these open weights yourself on rented GPUs. No API sells this model, so this is its price on the chart.
- Hardware
- 8× B200 180GB
- FP8 weights
- Throughput
- ~18,596 tok/s
- many requests batched
- Price
- $0.30–2.69 /M tok
- busy → light use
- On your own machine
- Mac Studio M5 Ultra 512GB (int4, ~42 tok/s single-stream)
- single consumer GPU or Mac
Assumptions (5)
- FP8 weights (744B params, 40B active per token) + 25% KV-cache headroom ≈ 930 GB
- 8× B200 180GB at $5.98–$14.24/GPU-hour on-demand (2026-09-24)
- ~18,596 output tok/s aggregate at batch 532 (bandwidth-bound); MoE compute scales with active params
- Blended 3:1 input:output; prefill ~212,625 tok/s
- Low = 75% utilization at the low GPU price; high = 20% at the high price
Benchmark scores
~ italic, dashed = no published score yet, estimated from related benchmarks.
Coding
SWE-bench VerifiedEstimated
~#3 of 14 (estimated)~81.9%
No published score — estimated from related benchmarks
SWE-bench Pro (Public)Lab-reported
#10 of 1859.6%
- #1Claude Opus 5.589.9%
- #2Claude Sonnet 5.581.3%
- #9Nex-N2.5-Pro61.2%
- #10Atria Dawn Preview59.6%
- #11MiniMax-M359.0%
huggingface.co · 2026-09-11
DeepSWE v1.1Estimated
~#27 of 31 (estimated)~54.2%
- #1DeepSeek V4.1 Flash74.2%
- #2Grok 4.772.6%
- #26Nex-N2.5-Pro55.8%
- ~#27Atria Dawn PreviewEstimated54.2%
- #27Qwen3.8 Max (0902)51.0%
No published score — estimated from related benchmarks
Terminal-Bench 4.0Estimated
~#18 of 84 (estimated)~19.9%
- #1Claude Sonnet 5.563.6%
- #2Claude Opus 5.559.6%
- #17MiMo-V2.6-Flash22.7%
- ~#18Atria Dawn PreviewEstimated19.9%
- #18Gemini 3.8 Flash19.7%
No published score — estimated from related benchmarks
LiveCodeBenchEstimated
~#6 of 20 (estimated)~81.3%
- #1Qwen3.8-Omni-Flash92.6%
- #2Qwen3.8 Flash-Next91.9%
- #5gpt-oss-120b87.8%
- ~#6Atria Dawn PreviewEstimated81.3%
- #6Gemma 4 31B80.0%
No published score — estimated from related benchmarks
SciCodeEstimated
~#12 of 96 (estimated)~57.3%
- #1Claude Opus 5.566.9%
- #2Claude Fable 5.163.1%
- #11Grok 4.757.4%
- ~#12Atria Dawn PreviewEstimated57.3%
- #12Gemini 3.8 Flash56.6%
No published score — estimated from related benchmarks
AA Coding Agent IndexEstimated
~#7 of 15 (estimated)~55.1
- #1Claude Opus 5.566.0
- #2Claude Fable 5.162.2
- #6Grok 4.756.3
- ~#7Atria Dawn PreviewEstimated55.1
- #7Muse Spark 1.354.3
No published score — estimated from related benchmarks
Agents
τ²-bench (Telecom)Estimated
~#26 of 116 (estimated)~82.5%
- #1Step 3.7 Flash98.5%
- #2Gemini 3.1 Pro95.6%
- #24Nemotron 3 Ultra83.3%
- ~#26Atria Dawn PreviewEstimated82.5%
- #26Nex-N2-Pro81.6%
No published score — estimated from related benchmarks
OSWorld-VerifiedEstimated
~#3 of 12 (estimated)~82.6%
No published score — estimated from related benchmarks
OSWorld 2.0Estimated
~#3 of 5 (estimated)~57.9%
No published score — estimated from related benchmarks
GDPval-AA (Elo)Estimated
~#18 of 87 (estimated)~1458 Elo
- #1Claude Opus 5.51846
- #2Claude Fable 5.11735
- #17GPT-6 Sol1487
- ~#18Atria Dawn PreviewEstimated1458
- #18Claude Sonnet 51449
No published score — estimated from related benchmarks
AutomationBench-AAEstimated
~#24 of 80 (estimated)~40.3%
- #1Claude Opus 5.569.5%
- #2DeepSeek V4.1 Flash68.9%
- #23GPT-5.547.3%
- ~#24Atria Dawn PreviewEstimated40.3%
- #24K2 Horizon 375B A23B37.2%
No published score — estimated from related benchmarks
τ-Bench Banking (AA)Estimated
~#21 of 89 (estimated)~32.3%
- #1Muse Spark 1.350.5%
- #2GLM-5.350.3%
- #20K2 Horizon 375B A23B34.2%
- ~#21Atria Dawn PreviewEstimated32.3%
- #21K2 Horizon MoVA 36B A4B32.0%
No published score — estimated from related benchmarks