DeepSeek V4.1 Flash
- Input
- $0.30/M tok
- Output
- $1.20/M tok
- Blended
- $0.525/M tok · 3:1
- Speed
- 232tok/s
- Context
- 1Mtokens
Open weights (MIT). Peak-hour price; off-peak is 50% off.
Rankings
Coding
Mixed#14 / 103rank by quality66.4±6.4quality / 1003/7 benchmarks measuredPrice frontierSpeed frontierWriting & chat
Verified#26 / 146rank by quality69.8±4.3quality / 1003/5 benchmarks measuredVision
Verified#35 / 82rank by quality56.4±4.8quality / 1003/6 benchmarks measuredHard reasoning
Mixed#23 / 182rank by quality76.1±4.0quality / 1003/4 benchmarks measuredAgents
Verified#13 / 173rank by quality82.1±7.8quality / 1002/7 benchmarks measuredSpeed frontier
Self-hostingEstimated
How we estimate →What it would cost to run these open weights yourself on rented GPUs, compared with the API price above.
- Hardware
- 8× B200 180GB
- FP8 weights
- Throughput
- ~39,657 tok/s
- many requests batched
- Price
- $0.14–1.22 /M tok
- busy → light use
- On your own machine
- Mac Studio M5 Ultra 512GB (int4, ~110 tok/s single-stream)
- single consumer GPU or Mac
Assumptions (5)
- FP8 weights (552B params, 16B active per token) + 40% KV-cache headroom ≈ 773 GB
- 8× B200 180GB at $5.98–$14.24/GPU-hour on-demand (2026-09-24)
- ~39,657 output tok/s aggregate at batch 1135 (bandwidth-bound); MoE compute scales with active params
- Blended 3:1 input:output; prefill ~531,563 tok/s
- Low = 75% utilization at the low GPU price; high = 20% at the high price
Benchmark scores
~ italic, dashed = no published score yet, estimated from related benchmarks.
Coding
SWE-bench VerifiedEstimated
~#3 of 14 (estimated)~83.2%
No published score — estimated from related benchmarks
SWE-bench Pro (Public)Estimated
~#6 of 19 (estimated)~65.9%
- #1Claude Opus 5.589.9%
- #2Claude Sonnet 5.581.3%
- #5Intern-S2-397B68.5%
- ~#6DeepSeek V4.1 FlashEstimated65.9%
- #6Hy4 preview65.7%
No published score — estimated from related benchmarks
Terminal-Bench 4.0Verified
#14 of 8326.8%
- #1Claude Sonnet 5.563.6%
- #2Claude Opus 5.559.6%
- #13GLM-5.3 Flash32.8%
- #14DeepSeek V4.1 Flash26.8%
- #15Grok 4.725.8%
artificialanalysis.ai · 2026-09-24
LiveCodeBenchEstimated
~#5 of 20 (estimated)~88.7%
- #1Qwen3.8-Omni-Flash92.6%
- #2Qwen3.8 Flash-Next91.9%
- #4Nemotron 3 Ultra89.0%
- ~#5DeepSeek V4.1 FlashEstimated88.7%
- #5gpt-oss-120b87.8%
No published score — estimated from related benchmarks
SciCodeVerified
#24 of 9551.9%
- #1Claude Opus 5.566.9%
- #2Claude Fable 5.163.1%
- #22Qwen3.8 Max (0902)52.1%
- #24DeepSeek V4.1 Flash51.9%
- #25GLM-5.3 Flash51.6%
artificialanalysis.ai · 2026-09-24
AA Coding Agent IndexEstimated
~#10 of 15 (estimated)~48.0
- #1Claude Opus 5.566.0
- #2Claude Fable 5.162.2
- #9Kimi K351.9
- ~#10DeepSeek V4.1 FlashEstimated48.0
- #10Grok 4.647.0
No published score — estimated from related benchmarks
Writing & chat
IFBenchEstimated
~#19 of 126 (estimated)~75.5%
- #1MiniMax-M382.9%
- #2Qwen3.8 2.4T A95B82.8%
- #18Gemma 4 31B75.6%
- ~#19DeepSeek V4.1 FlashEstimated75.5%
- #19GPT-5.3 Codex75.4%
No published score — estimated from related benchmarks
LMArena Text, Non-English (Elo)Verified
#12 of 531462 Elo
- #1Claude Opus 5.51501
- #2Claude Fable 5.1Few votes1496
- #11GLM-5.3 Flash1463
- #12DeepSeek V4.1 Flash1462
- #13MiMo-V2.6-Pro1461
lmarena.ai · 2026-09-25
LMArena Text (Elo)Verified
#13 of 531477 Elo
- #1Claude Opus 5.51509
- #2Claude Fable 5.11501
- #10MiMo-V2.6-Pro1480
- #13DeepSeek V4.1 Flash1477
- #14GLM-5.3 Flash1475
lmarena.ai · 2026-09-25
Vision
MMMU-ProVerified
#22 of 7677.0%
- #1Claude Opus 5.587.7%
- #2GPT-6 Astra86.9%
- #20Qwen3.5 397B A17B77.3%
- #22DeepSeek V4.1 Flash77.0%
- #23Intern-S2-Preview (35B-A3B)76.9%
artificialanalysis.ai · 2026-09-24
CharXiv ReasoningEstimated
~#5 of 9 (estimated)~79.2%
- #1Gemini 3.8 Flash86.2%
- #2Kimi K384.8%
- #4Qwen3.8-Omni-Flash83.5%
- ~#5DeepSeek V4.1 FlashEstimated79.2%
- #5Muse Glimmer78.8%
No published score — estimated from related benchmarks
OmniDocBench v1.5Estimated
~#5 of 8 (estimated)~85.2%
- #1MiniMax-M391.6%
- #2Kimi K391.1%
- #4Gemini 3.1 Pro85.3%
- ~#5DeepSeek V4.1 FlashEstimated85.2%
- #5Claude Haiku 4.579.6%
No published score — estimated from related benchmarks
GDP.pdf (AA)Verified
#24 of 4212.8%
- #1GPT-6 Astra31.0%
- #2Muse Spark 1.326.6%
- #23Claude Sonnet 513.2%
- #24DeepSeek V4.1 Flash12.8%
- #24Inkling12.8%
artificialanalysis.ai · 2026-09-24
ChartographyVerified
#19 of 2211.6%
- #1GPT-6 Astra71.0%
- #2Claude Opus 5.566.3%
- #18Grok 4.714.7%
- #19DeepSeek V4.1 Flash11.6%
- #20Gemini 3.5 Flash-Lite10.2%
surgehq.ai · 2026-09-24
Hard reasoning
GPQA DiamondLab-reported
#23 of 17290.9%
- #1GPT-6 Astra96.1%
- #2Gemini 3.8 Flash95.3%
- #22Qwen3.8-Omni-Flash91.0%
- #23DeepSeek V4.1 Flash90.9%
- #24Qwen3.8 27B90.5%
huggingface.co · 2026-09-10
Humanity's Last ExamVerified
#25 of 17839.2%
- #1Claude Opus 5.561.4%
- #2Claude Fable 5.159.1%
- #24GLM-5.3 Flash39.9%
- #25DeepSeek V4.1 Flash39.2%
- #26MiniMax-M339.0%
artificialanalysis.ai · 2026-09-24
AIME (latest)Estimated
~#2 of 25 (estimated)~96.7%
No published score — estimated from related benchmarks
Agents
τ²-bench (Telecom)Estimated
~#4 of 116 (estimated)~95.5%
- #1Step 3.7 Flash98.5%
- #2Gemini 3.1 Pro95.6%
- #2Qwen3.5 397B A17B95.6%
- ~#4DeepSeek V4.1 FlashEstimated95.5%
- #4Qwen3.6 35B A3B95.3%
No published score — estimated from related benchmarks
OSWorld-VerifiedEstimated
~#5 of 12 (estimated)~81.7%
- #1Kimi K384.8%
- #2Qwen3.8 27B84.3%
- #4MiMo-V2.6-Pro82.0%
- ~#5DeepSeek V4.1 FlashEstimated81.7%
- #5Claude Sonnet 581.2%
No published score — estimated from related benchmarks
OSWorld 2.0Estimated
~#3 of 5 (estimated)~57.7%
No published score — estimated from related benchmarks
BrowseCompEstimated
~#7 of 18 (estimated)~87.6%
- #1Atria Dawn Preview92.5%
- #2GPT-6 Astra91.5%
- #6Step 5 Preview88.7%
- ~#7DeepSeek V4.1 FlashEstimated87.6%
- #7GPT-5.6 Terra87.5%
No published score — estimated from related benchmarks
GDPval-AA (Elo)Verified
#12 of 861600 Elo
- #1Claude Opus 5.51846
- #2Claude Fable 5.11735
- #11Qwen3.8 Flash-Next1612
- #12DeepSeek V4.1 Flash1600
- #13Qwen3.8 2.4T A95B1597
artificialanalysis.ai · 2026-09-24
τ-Bench Banking (AA)Estimated
~#14 of 89 (estimated)~40.7%
- #1Muse Spark 1.350.5%
- #2GLM-5.350.3%
- #13GPT-6 Astra41.4%
- ~#14DeepSeek V4.1 FlashEstimated40.7%
- #14GPT-5.6 Terra40.2%
No published score — estimated from related benchmarks