Skip to main content
AI Choice Engine

DeepSeek

DeepSeek V4.1 Flash

DeepSeek V4.1 Flash is a DeepSeek model with 1M context, $0.30 in / $1.20 out per million tokens, last verified September 26, 2026. Open weights: yes. Public-board scores are linked to the publisher; vendor-only figures stay in the watch-out, not on the board.

DeepSeek V4.1 Flash is DeepSeek's new default — an open-weight (MIT) multimodal MoE that DeepSeek says beats V4-Pro on performance, cost and speed.

Catalog record checked September 26, 2026Individual provider fields may change

AI model specification details
SpecificationDeepSeek V4.1 Flash
ProviderDeepSeek
TierBalanced
Context window1M
Max output384K
Input / 1M tokens$0.30
Output / 1M tokens$1.20
WeightsOpen
Parameters552B MoE (causal encoder-decoder; 8B active for input, 16B for output)
Reasoning levelslow, high, max
Modalitiestext, image
LicenseMIT
API model iddeepseek-flash
ReleasedSeptember 10, 2026

Pricing tiers: Official peak rates $0.30/$1.20 per MTok; off-peak $0.15/$0.60 (50% of peak). Cache hit $0.006 peak / $0.003 off-peak. Peak hours are weekdays 01:00–04:00 and 06:00–10:00 UTC. Open weights (MIT). The legacy deepseek-v4-flash and deepseek-v4-flash-vision-exp ids temporarily route here.

Verified evidence

Published benchmark results

Each result keeps its source and measurement date visible. A missing benchmark is not treated as a zero.

Published benchmark results for DeepSeek V4.1 Flash
BenchmarkScoreMeasuredSource
Artificial Analysis Intelligence Index [max]39.52026-09-26Artificial Analysis · View source
DeepSWE 1.1 [max]74.22026-09-10DeepSeek V4.1-Flash release notes (vendor; not yet on Datacurve's board) · View source
Terminal-Bench 2.1 [max]90.62026-09-10DeepSeek V4.1-Flash release notes (vendor) · View source
GPQA Diamond [max]90.92026-09-10DeepSeek V4.1-Flash release notes (vendor) · View source

Best for

  • Open-weight price-performance
  • High-volume hosted agents
  • Self-hosting with MIT weights

Watch out

Vendor benchmark claims are far ahead of its independent Artificial Analysis score (39.5 on v4.3.2); peak/off-peak billing means the hour you run matters.

History

DeepSeek V4.1 Flash timeline

Release and status events for this model that passed two-source verification — dated, sourced, and linked.

  • DeepSeek V4.1-FlashReleaseSeptember 10, 2026

    DeepSeek releases V4.1-Flash (MIT weights, 1M context, image input) as its new default at $0.30/$1.20 per MTok peak ($0.15/$0.60 off-peak) under the API id `deepseek-flash`, and says it beats V4-Pro on performance, cost and speed.

    DeepSeek — V4.1-Flash release notes

Source receipts

Catalog figures for this model were checked against the following sources.

Common questions

DeepSeek V4.1 Flash

Answered from the verified figures on this page rather than general guidance.

How much does DeepSeek V4.1 Flash cost per million tokens?
DeepSeek V4.1 Flash is listed at $0.30 per million input tokens and $1.20 per million output tokens at standard rates. Output tokens usually dominate real bills, so weigh the output rate more heavily than the input rate. Official peak rates $0.30/$1.20 per MTok; off-peak $0.15/$0.60 (50% of peak). Cache hit $0.006 peak / $0.003 off-peak. Peak hours are weekdays 01:00–04:00 and 06:00–10:00 UTC. Open weights (MIT). The legacy deepseek-v4-flash and deepseek-v4-flash-vision-exp ids temporarily route here.
What is DeepSeek V4.1 Flash's context window?
DeepSeek V4.1 Flash accepts about 1M tokens of context. That only matters if you routinely send very long documents, large codebases, or multi-turn histories that approach that limit.
What is DeepSeek V4.1 Flash best for?
DeepSeek V4.1 Flash is a balanced tier from DeepSeek. It suits open-weight price-performance, high-volume hosted agents, self-hosting with mit weights. Vendor benchmark claims are far ahead of its independent Artificial Analysis score (39.5 on v4.3.2); peak/off-peak billing means the hour you run matters.
Can I self-host DeepSeek V4.1 Flash?
DeepSeek V4.1 Flash publishes open weights, but self-hosting depends on the licence, hardware footprint, quantisation quality, and serving stack. A hosted API is often cheaper until you have measured throughput and concurrency on your own hardware.

Compare it

Featured head-to-head comparisons

A deliberately selected pairing for this model, rather than an automatically generated matrix.