ModelPriceWatch.com
Last scan 2026-09-22 Models tracked 267 Providers 35 Cheapest paid Granite 4.0 H Micro $0.017/Mtok in Every price links to its source

DeepSeek V4 Pro

by DeepSeek

Current mid tier mid tier
Availability: Generally available, with no end-of-life date. DeepSeek REVERSED the retirement it had announced for this model. On 2026-09-10 footnote (2) of its rate card said V4.1 Flash had "comprehensively surpassed" V4 Pro and that from 12:00 Beijing Time on 2026-09-14 all deepseek-v4-pro requests would be routed to V4.1 Flash; we staged that flip. Re-read 2026-09-11T14:59Z, the SAME footnote (2) now reads: "In response to user demand, we have decided to continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged. We will provide further notice should there be any changes." All retirement language for this model is gone from the page. The reversal is on three of DeepSeek's own surfaces — the rate card, quick_start/first_api_call, and the /updates change log, whose 2026-09-10 entry was itself edited to carry it. So the staged flip was withdrawn before it could fire, and this row stays Current at the rates shown, which the same read re-confirmed unchanged (MODEL VERSION still DeepSeek-V4-Pro-0813). DeepSeek promises "further notice should there be any changes", so a future retirement is possible but is undated and unannounced; nothing is publishable from it today.

Today's price · per 1M tokens

Input

$1.32

per 1M tokens

Output

$3.96

per 1M tokens

Blended

$1.98

blended $/1M — 3:1 weighted input:output

Cached input

$0.044

3% of input — prompt caching

How this price is scoped: DeepSeek bills this model at TWO rates depending on the hour. The tracked figure is the PEAK rate: $1.32 input / $0.044 cached input / $3.96 output per 1M tokens. Off-peak is exactly half: $0.66 / $0.022 / $1.98. Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday; every other hour, and all of Saturday and Sunday, is off-peak. Peak is the headline here because DeepSeek defines off-peak as half of peak rather than the other way round, so peak is the list rate, and because a caller who does not schedule around the clock needs the published number to be a ceiling rather than a floor. Halve it for a workload that runs entirely outside those hours. Until 16:00 UTC on 2026-08-16 this model billed a single flat $0.435 / $0.003625 / $0.87. These rates were announced on 2026-09-10 to stop being purchasable at 2026-09-14T04:00Z; DeepSeek withdrew that on 2026-09-11, stating the billing method is unchanged past 2026-09-14, so they remain the live rates.

Source: official DeepSeek pricing · read MODELPRICEWATCH.COM · 2026-09-22

Price receipt

We read DeepSeek's own pricing page on and found deepseek-v4-pro listed with a price on the same column — that read is where the number above comes from. We keep the page text we read; its content hash is 6a19d87d91.

Available on 4 hosts

cheapest blended first · $ per 1M tokens

DeepSeek V4 Pro is sold by 4 providers. Prices are per 1M tokens (blended = (3×input + 1×output) ÷ 4). The first-party row is the model maker; “vs first-party” shows each host’s blended price relative to it.

DeepSeek V4 Pro price on every host still selling it, per 1M tokens, cheapest blended first
Host Input Output Blended vs first-party
DeepInfra details → $1.30 $2.60 $1.63 -18%
DeepSeek first-party $1.32 $3.96 $1.98
Fireworks details → $1.32 $3.96 $1.98 0%
Together details → $1.32 $3.96 $1.98 0%
Each host links to its provider page and official pricing. Prices refresh twice daily. MODELPRICEWATCH.COM · 2026-09-22

Overview

DeepSeek's open-weight (MIT) flagship — a 1.6T-parameter MoE (49B active) with a 1M-token context and thinking/non-thinking modes. Leads open models on world knowledge, math, and coding (~80.6% SWE-bench Verified, top open-weights). DeepSeek announced on 2026-09-10 that it would retire this model on 2026-09-14, then reversed that on 2026-09-11: it now says V4 Pro keeps its API service and its billing past that date. Served at $1.32/$3.96 per 1M tokens at peak (01:00-04:00 and 06:00-10:00 UTC, Monday to Friday); half that in every other hour.

Capabilities

struck through = not supported
Input 1/4
Text Image Audio Video
Output
Text
Features 3/9
Prompt caching Reasoning Coding Fast inference Long context Open weights Multimodal Web search Realtime

Benchmark performance

accuracy % · higher is better
Avg benchmark score73.5%
Perf per $/Mtok37.1
GPQA Diamond
90.1%
SWE-Bench Verified vendor
80.6%
LiveCodeBench
93.5%
MCP Atlas
73.6%
BrowseComp
83.4%
Humanity's Last Exam
48.2%
ARC-AGI 2
30%
AIME 2025
75%
MMMLU
82%
BFCL
72%
HumanEval
80.6%
MATH 500
80%
Independent composite scores
  • AA Intelligence Index: 36.3 (Artificial Analysis composite across reasoning, knowledge and coding evals)
  • AA-Omniscience Index: 0.8 (−100–100)
  • GDPval-AA v2: 1493.3 ELO (Real-world work tasks, human baseline = 1000)
  • LMSYS Chatbot Arena: 1463.4 ELO (Human preference)
How it stacks up
  • Ranks #29 of 84 benchmarked models by average score
  • Ranks #50 of 181 comparably-measured models by percentile score, across 15 independent measurements
  • Ranks #96 of 181 comparably-tested models by normalized performance per dollar
  • Strongest at LiveCodeBench — 93.5%, #1 of 26

174.9 tokens/sec output 1.2s latency to first token (TTFT)

Source: Vellum LLM Leaderboard · updated · See full rankings →

Specifications

Provider
DeepSeek
Context window
1M tokens
Modality
text
Parameters
Proprietary
Open source
No — proprietary
Released
Status
Current
Last updated
Tags
reasoningcachingmid-tier

Availability verified: listed on DeepSeek's own page