DeepSeek
DeepSeek V4-Flash
Open-weight price-performance pick: MIT weights, 1M context, and API rates far below closed budget tiers.
Catalog record checked August 14, 2026; individual provider fields may change.
| Specification | DeepSeek V4-Flash |
|---|---|
| Provider | DeepSeek |
| Tier | Budget |
| Context window | 1M |
| Max output | 384K |
| Input / 1M tokens | $0.14 |
| Output / 1M tokens | $0.28 |
| Weights | Open |
| Parameters | 284B total / 13B active (MoE) |
| Modalities | text |
| Released | July 31, 2026 |
Pricing tiers: Current list is $0.14/$0.28 per million tokens (cache miss / output) until 2026-08-16 16:00 UTC. From then, official peak/off-peak rates: off-peak $0.22/$0.66, peak $0.44/$1.32 (cache miss / output). Peak hours are 01:00–04:00 and 06:00–10:00 UTC.
Verified evidence
Published benchmark results
Each result keeps its source and measurement date visible. A missing benchmark is not treated as a zero.
Artificial Analysis Intelligence Index [max]
2026-08-08 · Artificial Analysis
Benchmark
DeepSWE 1.1 in context
The full local snapshot puts this model beside the wider field, including cost per completed task.
Local leader
Claude Opus 5 [max]
74%
Rows shown
24
Highest published reasoning effort per model (not best Pass@1)
Snapshot date
2026-08-13
Mirrored from deepswe.datacurve.ai
Better is toward the top-right (higher pass rate, lower cost). X-axis is reversed to match DeepSWE’s public chart. v1.1 uses average cost / tokens / steps; v1 uses published medians.
| # | Model | Pass@1 | Cost / task | Tokens / task | Steps / task |
|---|---|---|---|---|---|
| 1 | Claude Opus 5 [max] | 74% | $11.84 | 118k | 99 |
| 2 | GPT-5.6 Sol [max] | 73% | $8.39 | 60k | 61 |
| 3 | Claude Fable 5 [max] | 70% | $21.63 | 119k | 88 |
| 4 | GPT-5.6 Terra [max] | 70% | $4.95 | 72k | 76 |
| 5 | Kimi K3 [max] | 69% | $4.65 | 82k | 98 |
| 6 | GPT-5.6 Luna [max] | 67% | $3.03 | 73k | 102 |
| 7 | GPT-5.5 [xhigh] | 67% | $7.23 | 46k | 82 |
| 8 | Grok 4.6 [xhigh] | 67% | $5.50 | 71k | 87 |
| 9 | Gemini 3.7 Flash [high] | 65% | $2.18 | 107k | 125 |
| 10 | DeepSeek V4-Pro [max] | 63% | $0.24 | 106k | 155 |
| 11 | Claude Opus 4.8 [max] | 59% | $13.22 | 135k | 120 |
| 12 | Qwen3.8-Max [xhigh] | 58% | $3.73 | 95k | 111 |
| 13 | Muse Spark 1.2 [xhigh] | 55% | $3.70 | 99k | 101 |
| 14 | Claude Sonnet 5 [max] | 54% | $26.40 | 214k | 268 |
| 15 | Grok 4.5 [high] | 54% | $2.42 | 36k | 61 |
| 16 | DeepSeek V4-Flash [max] | 53% | $0.10 | 108k | 153 |
| 17 | Muse Spark 1.1 [xhigh] | 53% | $2.36 | 74k | 96 |
| 18 | GPT-5.4 [xhigh] | 52% | $5.65 | 71k | 70 |
| 19 | Gemini 3.6 Flash [high] | 47% | $4.42 | 96k | 117 |
| 20 | GLM 5.2 [max] | 44% | $3.92 | 78k | 129 |
| 21 | Gemini 3.5 Flash [high] | 36% | $3.45 | 76k | 105 |
| 22 | Kimi K2.7 Code | 31% | $2.82 | 59k | 149 |
| 23 | Claude Sonnet 4.6 [high] | 30% | $5.52 | 76k | 134 |
| 24 | Gemini 3.1 Pro [high] | 12% | $2.14 | 28k | 76 |
DeepSWE “Best” picks the highest published reasoning effort per model (not the highest pass rate). Small gaps may not be statistically meaningful — confirm on deepswe.datacurve.ai.
Best for
- Cost-sensitive hosted agents
- High-volume coding assist
- Open-weight deployments with cluster VRAM
Watch out
Common questions
DeepSeek V4-Flash
Answered from the verified figures on this page rather than general guidance.
How much does DeepSeek V4-Flash cost per million tokens?
DeepSeek V4-Flash is listed at $0.14 per million input tokens and $0.28 per million output tokens at standard rates. Output tokens usually dominate real bills, so weigh the output rate more heavily than the input rate. Current list is $0.14/$0.28 per million tokens (cache miss / output) until 2026-08-16 16:00 UTC. From then, official peak/off-peak rates: off-peak $0.22/$0.66, peak $0.44/$1.32 (cache miss / output). Peak hours are 01:00–04:00 and 06:00–10:00 UTC.
What is DeepSeek V4-Flash's context window?
DeepSeek V4-Flash accepts about 1M tokens of context. That only matters if you routinely send very long documents, large codebases, or multi-turn histories that approach that limit.
What is DeepSeek V4-Flash best for?
DeepSeek V4-Flash is a budget tier from DeepSeek. It suits cost-sensitive hosted agents, high-volume coding assist, open-weight deployments with cluster vram. API list prices move to peak/off-peak from 2026-08-16 16:00 UTC (see pricing note). Self-hosting the ~284B MoE still needs roughly 90–170GB+ class VRAM depending on quant — not a laptop budget build. DeepSWE snapshot also uses the shared mini-swe-agent harness — verify cost, effort, and serving conditions before treating the result as a forecast.
Can I self-host DeepSeek V4-Flash?
DeepSeek V4-Flash publishes open weights, but self-hosting depends on the licence, hardware footprint, quantisation quality, and serving stack. A hosted API is often cheaper until you have measured throughput and concurrency on your own hardware.
Compare it
Head-to-head model comparisons
These are the published pairings that put this model against a plausible alternative.