Skip to main content

DeepSeek

DeepSeek V4-Flash

Open-weight price-performance pick: MIT weights, 1M context, and API rates far below closed budget tiers.

Catalog record checked August 14, 2026; individual provider fields may change.

AI model specification details
SpecificationDeepSeek V4-Flash
ProviderDeepSeek
TierBudget
Context window1M
Max output384K
Input / 1M tokens$0.14
Output / 1M tokens$0.28
WeightsOpen
Parameters284B total / 13B active (MoE)
Modalitiestext
ReleasedJuly 31, 2026

Pricing tiers: Current list is $0.14/$0.28 per million tokens (cache miss / output) until 2026-08-16 16:00 UTC. From then, official peak/off-peak rates: off-peak $0.22/$0.66, peak $0.44/$1.32 (cache miss / output). Peak hours are 01:00–04:00 and 06:00–10:00 UTC.

Verified evidence

Published benchmark results

Each result keeps its source and measurement date visible. A missing benchmark is not treated as a zero.

Artificial Analysis Intelligence Index [max]

2026-08-08 · Artificial Analysis

Benchmark

DeepSWE 1.1 in context

The full local snapshot puts this model beside the wider field, including cost per completed task.

Local leader

Claude Opus 5 [max]

74%

Rows shown

24

Highest published reasoning effort per model (not best Pass@1)

Snapshot date

2026-08-13

Mirrored from deepswe.datacurve.ai

DeepSWE 1.1 pass@10%16%32%48%64%80%$0$4.50$9.00$13.50$18.00$22.50$27.00Avg cost per taskClaude Opus 5 [max]: 74% · $11.84 · 118k tokens · 99 stepsClaude Opus 5GPT-5.6 Sol [max]: 73% · $8.39 · 60k tokens · 61 stepsGPT-5.6 SolClaude Fable 5 [max]: 70% · $21.63 · 119k tokens · 88 stepsClaude Fable 5GPT-5.6 Terra [max]: 70% · $4.95 · 72k tokens · 76 stepsGPT-5.6 TerraKimi K3 [max]: 69% · $4.65 · 82k tokens · 98 stepsKimi K3GPT-5.6 Luna [max]: 67% · $3.03 · 73k tokens · 102 stepsGPT-5.6 LunaGPT-5.5 [xhigh]: 67% · $7.23 · 46k tokens · 82 stepsGPT-5.5Grok 4.6 [xhigh]: 67% · $5.50 · 71k tokens · 87 stepsGrok 4.6Gemini 3.7 Flash [high]: 65% · $2.18 · 107k tokens · 125 stepsGemini 3.7 FlashDeepSeek V4-Pro [max]: 63% · $0.24 · 106k tokens · 155 stepsDeepSeek V4-ProClaude Opus 4.8 [max]: 59% · $13.22 · 135k tokens · 120 stepsClaude Opus 4.8Qwen3.8-Max [xhigh]: 58% · $3.73 · 95k tokens · 111 stepsQwen3.8-MaxMuse Spark 1.2 [xhigh]: 55% · $3.70 · 99k tokens · 101 stepsMuse Spark 1.2Claude Sonnet 5 [max]: 54% · $26.40 · 214k tokens · 268 stepsClaude Sonnet 5Grok 4.5 [high]: 54% · $2.42 · 36k tokens · 61 stepsGrok 4.5DeepSeek V4-Flash [max]: 53% · $0.10 · 108k tokens · 153 stepsDeepSeek V4-FlashMuse Spark 1.1 [xhigh]: 53% · $2.36 · 74k tokens · 96 stepsMuse Spark 1.1GPT-5.4 [xhigh]: 52% · $5.65 · 71k tokens · 70 stepsGPT-5.4Gemini 3.6 Flash [high]: 47% · $4.42 · 96k tokens · 117 stepsGemini 3.6 FlashGLM 5.2 [max]: 44% · $3.92 · 78k tokens · 129 stepsGLM 5.2Gemini 3.5 Flash [high]: 36% · $3.45 · 76k tokens · 105 stepsGemini 3.5 FlashKimi K2.7 Code: 31% · $2.82 · 59k tokens · 149 stepsKimi K2.7 CodeClaude Sonnet 4.6 [high]: 30% · $5.52 · 76k tokens · 134 stepsClaude Sonnet 4.6Gemini 3.1 Pro [high]: 12% · $2.14 · 28k tokens · 76 stepsGemini 3.1 Pro

Better is toward the top-right (higher pass rate, lower cost). X-axis is reversed to match DeepSWE’s public chart. v1.1 uses average cost / tokens / steps; v1 uses published medians.

DeepSWE 1.1 leaderboard with pass rate, cost, tokens, and steps per task
#ModelPass@1Cost / taskTokens / taskSteps / task
1Claude Opus 5 [max]74%$11.84118k99
2GPT-5.6 Sol [max]73%$8.3960k61
3Claude Fable 5 [max]70%$21.63119k88
4GPT-5.6 Terra [max]70%$4.9572k76
5Kimi K3 [max]69%$4.6582k98
6GPT-5.6 Luna [max]67%$3.0373k102
7GPT-5.5 [xhigh]67%$7.2346k82
8Grok 4.6 [xhigh]67%$5.5071k87
9Gemini 3.7 Flash [high]65%$2.18107k125
10DeepSeek V4-Pro [max]63%$0.24106k155
11Claude Opus 4.8 [max]59%$13.22135k120
12Qwen3.8-Max [xhigh]58%$3.7395k111
13Muse Spark 1.2 [xhigh]55%$3.7099k101
14Claude Sonnet 5 [max]54%$26.40214k268
15Grok 4.5 [high]54%$2.4236k61
16DeepSeek V4-Flash [max]53%$0.10108k153
17Muse Spark 1.1 [xhigh]53%$2.3674k96
18GPT-5.4 [xhigh]52%$5.6571k70
19Gemini 3.6 Flash [high]47%$4.4296k117
20GLM 5.2 [max]44%$3.9278k129
21Gemini 3.5 Flash [high]36%$3.4576k105
22Kimi K2.7 Code31%$2.8259k149
23Claude Sonnet 4.6 [high]30%$5.5276k134
24Gemini 3.1 Pro [high]12%$2.1428k76

DeepSWE “Best” picks the highest published reasoning effort per model (not the highest pass rate). Small gaps may not be statistically meaningful — confirm on deepswe.datacurve.ai.

Best for

  • Cost-sensitive hosted agents
  • High-volume coding assist
  • Open-weight deployments with cluster VRAM

Watch out

API list prices move to peak/off-peak from 2026-08-16 16:00 UTC (see pricing note). Self-hosting the ~284B MoE still needs roughly 90–170GB+ class VRAM depending on quant — not a laptop budget build. DeepSWE snapshot also uses the shared mini-swe-agent harness — verify cost, effort, and serving conditions before treating the result as a forecast.

Common questions

DeepSeek V4-Flash

Answered from the verified figures on this page rather than general guidance.

How much does DeepSeek V4-Flash cost per million tokens?

DeepSeek V4-Flash is listed at $0.14 per million input tokens and $0.28 per million output tokens at standard rates. Output tokens usually dominate real bills, so weigh the output rate more heavily than the input rate. Current list is $0.14/$0.28 per million tokens (cache miss / output) until 2026-08-16 16:00 UTC. From then, official peak/off-peak rates: off-peak $0.22/$0.66, peak $0.44/$1.32 (cache miss / output). Peak hours are 01:00–04:00 and 06:00–10:00 UTC.

What is DeepSeek V4-Flash's context window?

DeepSeek V4-Flash accepts about 1M tokens of context. That only matters if you routinely send very long documents, large codebases, or multi-turn histories that approach that limit.

What is DeepSeek V4-Flash best for?

DeepSeek V4-Flash is a budget tier from DeepSeek. It suits cost-sensitive hosted agents, high-volume coding assist, open-weight deployments with cluster vram. API list prices move to peak/off-peak from 2026-08-16 16:00 UTC (see pricing note). Self-hosting the ~284B MoE still needs roughly 90–170GB+ class VRAM depending on quant — not a laptop budget build. DeepSWE snapshot also uses the shared mini-swe-agent harness — verify cost, effort, and serving conditions before treating the result as a forecast.

Can I self-host DeepSeek V4-Flash?

DeepSeek V4-Flash publishes open weights, but self-hosting depends on the licence, hardware footprint, quantisation quality, and serving stack. A hosted API is often cheaper until you have measured throughput and concurrency on your own hardware.

Compare it

Head-to-head model comparisons

These are the published pairings that put this model against a plausible alternative.