DeepSeek · Model release
DeepSeek V4.1-Flash: the new default that retired V4-Flash
DeepSeek's September release is an MIT-licensed multimodal MoE that DeepSeek says beats V4-Pro on performance, cost and speed — and it retired V4-Flash on the same day.
What changed
- Released 2026-09-10 with MIT open weights, a 1M-token context, 384K max output and image input, on the new `deepseek-flash` API id.
- V4-Flash and V4-Flash-Vision-Exp were retired the same day; their legacy ids temporarily route to V4.1-Flash and bill at its rates.
- Priced $0.30/$1.20 per MTok at peak and $0.15/$0.60 off-peak (cache hits $0.006/$0.003) — below V4-Flash's old $0.44/$1.32 peak list.
Best fit
- Cost-sensitive agents
- Open-weight deployments
- High-volume coding
What to be careful about
- Peak hours are 01:00–04:00 and 06:00–10:00 UTC on weekdays — schedule batch work off-peak.
- DeepSeek's benchmark claims (Terminal-Bench 2.1 90.6, DeepSWE 1.1 74.2 at max effort) are vendor-run; on the independent AA Intelligence Index v4.3.2 it scores 39.5 against 36.0 for V4-Pro 0813.
- Open weights still need datacentre-class memory for the full expert set.
Dataset snapshot: DeepSeek V4.1 Flash is also represented in the checked comparison dataset. Prices and benchmark figures carry their own field dates.