Decision guide
DeepSeek V4-Flash is retired: moving to V4.1-Flash
DeepSeek retired V4-Flash on 2026-09-10. The legacy id temporarily routes to V4.1-Flash, so you are probably already running the new model — this guide covers making that explicit.
Switch if
- You call `deepseek-v4-flash` or `deepseek-v4-flash-vision-exp` on the first-party API — those ids are retired and only temporarily re-routed.
- You want image input, which V4.1-Flash adds on the `deepseek-flash` id.
- You want the lower V4.1-Flash rates: $0.30/$1.20 peak and $0.15/$0.60 off-peak versus V4-Flash's old $0.44/$1.32 peak.
Stay if
- You self-host the V4-Flash-0731 MIT weights and have validated them — the weights remain usable offline.
- A frozen evaluation baseline requires the old checkpoint and you can serve it yourself.
- You need time to re-validate; the temporary re-route already serves V4.1-Flash, so plan the explicit cutover rather than assuming the old model.
Check before production traffic moves
- 01Switch the API id to `deepseek-flash` explicitly rather than relying on the temporary alias.
- 02Re-run structured-output and tool-call tests: outputs now come from a different model.
- 03Re-check peak/off-peak scheduling (weekday peaks 01:00–04:00 and 06:00–10:00 UTC).
This is a forced migration, not an optional upgrade: the first-party V4-Flash endpoint is retired and its id is only a temporary alias. The main work is making the model change visible in logs and evaluations, because the alias may already have changed behaviour under you.
Run a controlled trial first
A second option in a labelled pilot is cheaper than a production cutover. Score the trial, then migrate or roll back on evidence.
Pilot plan
- Diff recent outputs before and after 2026-09-10 on the same prompts to see whether the re-route changed quality.
- Run the frozen prompt set against `deepseek-flash` with your validators.
- Price the new rates against your peak/off-peak traffic mix.
Score the trial
- Validator pass rate
- Accepted outputs versus the pre-2026-09-10 baseline.
- Cost per success
- Peak and off-peak spend per accepted task.
- Alias exposure
- Share of traffic still calling the retired id.
Migration sequence
- Replace the retired ids with `deepseek-flash` in configuration and deploy behind a flag.
- Alert on any remaining calls to `deepseek-v4-flash` so the temporary alias cannot silently disappear under you.
Rollback plan
- There is no first-party rollback to V4-Flash; the fallback is another provider or self-hosted V4-Flash weights.
- Keep a second-provider route configured before the temporary alias is withdrawn.
Hidden costs to price
- Self-hosting the old weights to avoid the change moves cost to GPUs and operations.
- Re-validation time for a model you did not choose to change.
Evidence to check