Best LLM APIs with Prompt Caching
LLM APIs that support prompt caching, ranked by cached-input price. Cache hits cost 50–99% less than fresh input — the biggest lever for cutting cost on repeated system prompts, RAG context, and long conversations.
MODELPRICEWATCH.COM · 2026-08-15
Cost calculator for this use casemonthly cost, top 3 models
- 1 DeepSeek V4 Flash
- $—
- 2 DeepSeek V4 Pro
- $—
- 3 Solar Pro 4
- $—
Full ranking — top 20 models
list prices, USD per 1M tokens| Model | Cached input | Blended* | Input | Output | Context | Provider |
|---|---|---|---|---|---|---|
| 1DeepSeek V4 Flash | $0.003 | $0.175 | $0.140 | $0.280 | 1M | DeepSeek |
| 2DeepSeek V4 Pro | $0.004 | $0.544 | $0.435 | $0.870 | 1M | DeepSeek |
| 3Solar Pro 4 | $0.006 | $0.052 | $0.030 | $0.120 | 524K | Upstage |
| 4GLM-4.7-FlashX | $0.010 | $0.153 | $0.070 | $0.400 | 128K | Z.AI |
| 5GPT-OSS 120B | $0.015 | $0.262 | $0.150 | $0.600 | 128K | Fireworks |
| 6Solar Pro 3 | $0.015 | $0.262 | $0.150 | $0.600 | 131K | Upstage |
| 7GLM-4.5-Air | $0.030 | $0.425 | $0.200 | $1.10 | 128K | Z.AI |
| 8KAT-Coder-Air V2.5 | $0.030 | $0.259 | $0.148 | $0.593 | 256K | Kwaipilot |
| 9GPT-OSS 20B | $0.035 | $0.128 | $0.070 | $0.300 | 128K | Fireworks |
| 10GLM-4.6V | $0.050 | $0.450 | $0.300 | $0.900 | 128K | Z.AI |
| 11MiniMax-M2.7 | $0.060 | $0.525 | $0.300 | $1.20 | 205K | MiniMax |
| 12MiniMax-M3 | $0.060 | $0.525 | $0.300 | $1.20 | 1M | MiniMax |
| 13Inkling Small | $0.060 | $0.525 | $0.300 | $1.20 | 262K | Thinking Machines |
| 14Qwen3.7-Plus | $0.080 | $0.700 | $0.400 | $1.60 | 128K | Fireworks |
| 15Kimi K2.5 | $0.100 | $1.20 | $0.600 | $3.00 | 262K | Moonshot |
| 16GLM-4.5 | $0.110 | $1.00 | $0.600 | $2.20 | 128K | Z.AI |
| 17GLM-4.6 | $0.110 | $1.00 | $0.600 | $2.20 | 200K | Z.AI |
| 18GLM-4.7 | $0.110 | $1.00 | $0.600 | $2.20 | 128K | Z.AI |
| 19NVIDIA Nemotron 3 Ultra | $0.120 | $1.05 | $0.600 | $2.40 | 128K | Fireworks |
| 20Qwen3.7-Max | $0.130 | $1.88 | $1.25 | $3.75 | 1M | Together |
Recent price movement in this ranking
price-only deltas · logged by the daily scan1 of the top 20 prompt caching models has re-priced since we began tracking · last checked Aug 15, 2026.
| Model | Changes | Latest move | Provider |
|---|---|---|---|
| MiniMax-M3 | 1 | price cut on Jun 17, 2026: $0.6/$2.4 → $0.3/$1.2 /Mtok | MiniMax |
Most recently, MiniMax-M3 cut its price on Jun 17, 2026 — a sign pricing in this category is still moving, so re-check before committing to a long-term choice. See all recent price moves →
How models are selected
Models offering prompt caching, sorted by cached-input price per million tokens (cheapest cache reads first).
Prices are per million tokens (Mtok), sourced directly from — and linked to — official provider pricing pages. "Blended cost" is (3×input + 1×output) ÷ 4 — weighted toward input because real workloads read far more tokens than they generate.