ModelPriceWatch.com
Last scan 2026-09-22 Models tracked 267 Providers 35 Cheapest paid Granite 4.0 H Micro $0.017/Mtok in Every price links to its source

DeepSeek V4 Flash

by DeepSeek

Retired budget cheap tier
Retired on . No longer served on the endpoint we price — the figures below are kept for historical reference only. Replaced by DeepSeek V4.1 Flash. Full retirement status & replacements →
Availability: Retired 2026-09-10. DeepSeek's own rate card says it in footnote (1): "The legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp are still accepted, but the corresponding models have been retired, their requests are served by the DeepSeek-V4.1-Flash model and billed at the Flash price." The same-day change log entry repeats it. So the slug still resolves and still bills, but neither the model behind it nor the price is this row's — a call routed today costs the V4.1 Flash rate ($0.30 input / $0.006 cached input / $1.20 output per 1M at peak, half that off-peak), not the $0.44/$0.014/$1.32 archived here. Successor: DeepSeek V4.1 Flash (deepseek-v4-1-flash). Third-party hosts still serve their own copies of V4 Flash weights at their own prices; those rows are unaffected by DeepSeek's first-party retirement.

Today's price · per 1M tokens

Input

$0.440

per 1M tokens

Output

$1.32

per 1M tokens

Blended

$0.660

blended $/1M — 3:1 weighted input:output

Cached input

$0.014

3% of input — prompt caching

How this price is scoped: ARCHIVED first-party rate — what this model last cost, not what a call costs today. DeepSeek billed it at TWO rates depending on the hour. The tracked figure is the PEAK rate: $0.44 input / $0.014 cached input / $1.32 output per 1M tokens. Off-peak was exactly half: $0.22 / $0.007 / $0.66. Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday; every other hour is off-peak. Until 16:00 UTC on 2026-08-16 this model billed a single flat $0.14 / $0.0028 / $0.28. On 2026-09-10 DeepSeek retired the model; the name now routes to DeepSeek V4.1 Flash and bills at that model's price.

Source: official DeepSeek pricing · read MODELPRICEWATCH.COM · 2026-09-22

Price receipt

We read DeepSeek's own pricing page on and found deepseek-v4-flash OFF-PEAK $0.007 $0.22 $0.66 PEAK $0.014 $0 listed with a price on the same row — that read is where the number above comes from, recorded as a price increase. We keep the page text we read; its content hash is 2e222e11da.

Available on 2 hosts

cheapest blended first · $ per 1M tokens

DeepSeek V4 Flash is sold by 2 providers. Prices are per 1M tokens (blended = (3×input + 1×output) ÷ 4). “vs cheapest” compares each host to the lowest blended price.

DeepSeek V4 Flash price on every host still selling it, per 1M tokens, cheapest blended first
Host Input Output Blended vs cheapest
DeepInfra first-party details → $0.090 $0.180 $0.113
Fireworks details → $0.140 $0.280 $0.175 +56%
Each host links to its provider page and official pricing. Prices refresh twice daily. MODELPRICEWATCH.COM · 2026-09-22

Overview

Retired by DeepSeek on 2026-09-10, the day DeepSeek V4.1 Flash shipped. The deepseek-v4-flash model name is still accepted, but requests to it are served by DeepSeek V4.1 Flash and billed at the V4.1 Flash price, so any figure attributed to V4 Flash is what it last cost, not what you pay today. Last listed at $0.44/$1.32 per 1M tokens at peak.

Capabilities

struck through = not supported
Input 1/4
Text Image Audio Video
Output
Text
Features 4/9
Prompt caching Reasoning Coding Fast inference Long context Open weights Multimodal Web search Realtime

Benchmark performance

accuracy % · higher is better
Avg benchmark score77.5%
Perf per $/Mtok117.4
GPQA Diamond
88.1%
SWE-Bench Verified vendor
79%
LiveCodeBench
91.6%
MCP Atlas
69%
BrowseComp
85.9%
Humanity's Last Exam
51.6%
ARC-AGI 2 vendor
18%
AIME 2025 vendor
58%
MMMLU vendor
75%
BFCL vendor
62%
HumanEval
79%
MATH 500 vendor
68%
Independent composite scores
  • AA Intelligence Index: 34.5 (Artificial Analysis composite across reasoning, knowledge and coding evals)
  • AA-Omniscience Index: -14.3 (−100–100)
  • GDPval-AA v2: 1558.4 ELO (Real-world work tasks, human baseline = 1000)
  • LMSYS Chatbot Arena: 1435.6 ELO (Human preference)
How it stacks up
  • Ranks #16 of 84 benchmarked models by average score
  • Ranks #65 of 181 comparably-measured models by percentile score, across 10 independent measurements
  • Ranks #29 of 181 comparably-tested models by normalized performance per dollar
  • Strongest at LiveCodeBench — 91.6%, #5 of 26

107.9 tokens/sec output 1.42s latency to first token (TTFT)

Source: Vellum LLM Leaderboard · updated · See full rankings →

Specifications

Provider
DeepSeek
Context window
1M tokens
Modality
text
Parameters
Proprietary
Open source
No — proprietary
Released
Status
Retired ended
Last updated
Tags
fastreasoningcachingbudgetretired

Availability verified: per DeepSeek's own deprecation notice