ModelPriceWatch.com
Last scan 2026-09-11 Models tracked 258 Providers 32 Cheapest paid Granite 4.0 H Micro $0.017/Mtok in Every price links to its source

DeepSeek V4.1 Flash

by DeepSeek

Current new · 1d budget cheap tier
Availability: Call it as deepseek-flash — that is the model name on DeepSeek's own rate card, where the MODEL VERSION column reads DeepSeek-V4.1-Flash. The retired names deepseek-v4-flash and deepseek-v4-flash-vision-exp are still accepted and route here, billed at this model's price. From 2026-09-14T04:00Z (12:00 Beijing) deepseek-v4-pro routes here too, until DeepSeek releases V4.1 Pro.

Today's price · per 1M tokens

Input

$0.300

per 1M tokens

Output

$1.20

per 1M tokens

Blended

$0.525

blended $/1M — 3:1 weighted input:output

Cached input

$0.006

2% of input — prompt caching

How this price is scoped: DeepSeek bills this model at TWO rates depending on the hour. The tracked figure is the PEAK rate: $0.30 input (cache miss) / $0.006 input (cache hit) / $1.20 output per 1M tokens. Off-peak is exactly half: $0.15 / $0.003 / $0.60. Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday; every other hour, and all of Saturday and Sunday, is off-peak. Peak is the headline here because DeepSeek defines off-peak as half of peak rather than the other way round, so peak is the list rate, and because a caller who does not schedule around the clock needs the published number to be a ceiling rather than a floor. Halve it for a workload that runs entirely outside those hours. Third-party catalogues that quote this model at $0.15/$0.60 per 1M are printing the OFF-PEAK tier, not a different rate card.

Source: official DeepSeek pricing · read MODELPRICEWATCH.COM · 2026-09-11

Price receipt

We read DeepSeek's own pricing page on and found DeepSeek-V4.1-Flash listed with a price on the same column — that read is where the number above comes from, recorded as a new model. We keep the page text we read; its content hash is 34e56fb4e1.

Overview

DeepSeek's first model on its new causal encoder-decoder architecture — a 552B-parameter MoE activating 8B parameters for input and 16B for output, with native visual understanding, a 1M-token context and 384K max output. Called as deepseek-flash, it replaces both V4 Flash and V4 Flash Vision Exp. $0.30/$1.20 per 1M tokens at peak (01:00-04:00 and 06:00-10:00 UTC, Monday to Friday); half that in every other hour.

Capabilities

struck through = not supported
Input 2/5
Text Image Audio Video PDF
Output 1/5
Text Image Audio Video Embedding
Features 5/9
Prompt caching Reasoning Coding Fast inference Long context Open weights Multimodal Web search Realtime

Specifications

Provider
DeepSeek
Context window
1M tokens
Modality
text, image
Parameters
Proprietary
Open source
No — proprietary
Released
Status
Current
Last updated
Tags
fastreasoningmultimodalcachingbudget

Availability verified: listed on DeepSeek's own page