Today's price · per 1M tokens
Input
$0.440
per 1M tokens
Output
$1.32
per 1M tokens
Blended
$0.660
blended $/1M — 3:1 weighted input:output
Cached input
$0.014
3% of input — prompt caching
How this price is scoped: ARCHIVED first-party rate — what this model last cost, not what a call costs today. DeepSeek billed it at TWO rates depending on the hour. The tracked figure is the PEAK rate: $0.44 input / $0.014 cached input / $1.32 output per 1M tokens. Off-peak was exactly half: $0.22 / $0.007 / $0.66. Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday; every other hour is off-peak. Until 16:00 UTC on 2026-08-16 this model billed a single flat $0.14 / $0.0028 / $0.28. On 2026-09-10 DeepSeek retired the model; the name now routes to DeepSeek V4.1 Flash and bills at that model's price.
Price receipt
We read DeepSeek's own pricing page on and found deepseek-v4-flash OFF-PEAK $0.007 $0.22 $0.66 PEAK $0.014 $0 listed with a price on the same row — that read is where the number above comes from, recorded as a price increase. We keep the page text we read; its content hash is 2e222e11da.
Available on 2 hosts
cheapest blended first · $ per 1M tokensDeepSeek V4 Flash is sold by 2 providers. Prices are per 1M tokens (blended = (3×input + 1×output) ÷ 4). “vs cheapest” compares each host to the lowest blended price.
| Host | Input | Output | Blended | vs cheapest |
|---|---|---|---|---|
| DeepInfra first-party details → | $0.090 | $0.180 | $0.113 | — |
| Fireworks details → | $0.140 | $0.280 | $0.175 | +56% |
Overview
Retired by DeepSeek on 2026-09-10, the day DeepSeek V4.1 Flash shipped. The deepseek-v4-flash model name is still accepted, but requests to it are served by DeepSeek V4.1 Flash and billed at the V4.1 Flash price, so any figure attributed to V4 Flash is what it last cost, not what you pay today. Last listed at $0.44/$1.32 per 1M tokens at peak.
Capabilities
struck through = not supportedBenchmark performance
accuracy % · higher is better- AA Intelligence Index: 34.5 (Artificial Analysis composite across reasoning, knowledge and coding evals)
- AA-Omniscience Index: -14.3 (−100–100)
- GDPval-AA v2: 1558.4 ELO (Real-world work tasks, human baseline = 1000)
- LMSYS Chatbot Arena: 1435.6 ELO (Human preference)
- Ranks #16 of 84 benchmarked models by average score
- Ranks #65 of 181 comparably-measured models by percentile score, across 10 independent measurements
- Ranks #29 of 181 comparably-tested models by normalized performance per dollar
- Strongest at LiveCodeBench — 91.6%, #5 of 26
107.9 tokens/sec output 1.42s latency to first token (TTFT)
Source: Vellum LLM Leaderboard · updated · See full rankings →
Specifications
- Provider
- DeepSeek
- Context window
- 1M tokens
- Modality
- text
- Parameters
- Proprietary
- Open source
- No — proprietary
- Released
- Status
- Retired ended
- Last updated
- Tags
Availability verified: — per DeepSeek's own deprecation notice