Today's price · per 1M tokens
Input
$1.32
per 1M tokens
Output
$3.96
per 1M tokens
Blended
$1.98
blended $/1M — 3:1 weighted input:output
Cached input
$0.044
3% of input — prompt caching
How this price is scoped: DeepSeek bills this model at TWO rates depending on the hour. The tracked figure is the PEAK rate: $1.32 input / $0.044 cached input / $3.96 output per 1M tokens. Off-peak is exactly half: $0.66 / $0.022 / $1.98. Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday; every other hour, and all of Saturday and Sunday, is off-peak. Peak is the headline here because DeepSeek defines off-peak as half of peak rather than the other way round, so peak is the list rate, and because a caller who does not schedule around the clock needs the published number to be a ceiling rather than a floor. Halve it for a workload that runs entirely outside those hours. Until 16:00 UTC on 2026-08-16 this model billed a single flat $0.435 / $0.003625 / $0.87. These rates were announced on 2026-09-10 to stop being purchasable at 2026-09-14T04:00Z; DeepSeek withdrew that on 2026-09-11, stating the billing method is unchanged past 2026-09-14, so they remain the live rates.
Price receipt
We read DeepSeek's own pricing page on and found deepseek-v4-pro listed with a price on the same column — that read is where the number above comes from. We keep the page text we read; its content hash is 6a19d87d91.
Available on 4 hosts
cheapest blended first · $ per 1M tokensDeepSeek V4 Pro is sold by 4 providers. Prices are per 1M tokens (blended = (3×input + 1×output) ÷ 4). The first-party row is the model maker; “vs first-party” shows each host’s blended price relative to it.
| Host | Input | Output | Blended | vs first-party |
|---|---|---|---|---|
| DeepInfra details → | $1.30 | $2.60 | $1.63 | -18% |
| DeepSeek first-party | $1.32 | $3.96 | $1.98 | — |
| Fireworks details → | $1.32 | $3.96 | $1.98 | 0% |
| Together details → | $1.32 | $3.96 | $1.98 | 0% |
Overview
DeepSeek's open-weight (MIT) flagship — a 1.6T-parameter MoE (49B active) with a 1M-token context and thinking/non-thinking modes. Leads open models on world knowledge, math, and coding (~80.6% SWE-bench Verified, top open-weights). DeepSeek announced on 2026-09-10 that it would retire this model on 2026-09-14, then reversed that on 2026-09-11: it now says V4 Pro keeps its API service and its billing past that date. Served at $1.32/$3.96 per 1M tokens at peak (01:00-04:00 and 06:00-10:00 UTC, Monday to Friday); half that in every other hour.
Capabilities
struck through = not supportedBenchmark performance
accuracy % · higher is better- AA Intelligence Index: 36.3 (Artificial Analysis composite across reasoning, knowledge and coding evals)
- AA-Omniscience Index: 0.8 (−100–100)
- GDPval-AA v2: 1493.3 ELO (Real-world work tasks, human baseline = 1000)
- LMSYS Chatbot Arena: 1463.4 ELO (Human preference)
- Ranks #29 of 84 benchmarked models by average score
- Ranks #50 of 181 comparably-measured models by percentile score, across 15 independent measurements
- Ranks #96 of 181 comparably-tested models by normalized performance per dollar
- Strongest at LiveCodeBench — 93.5%, #1 of 26
174.9 tokens/sec output 1.2s latency to first token (TTFT)
Source: Vellum LLM Leaderboard · updated · See full rankings →
Specifications
- Provider
- DeepSeek
- Context window
- 1M tokens
- Modality
- text
- Parameters
- Proprietary
- Open source
- No — proprietary
- Released
- Status
- Current
- Last updated
- Tags
Availability verified: — listed on DeepSeek's own page