Best LLM APIs with Prompt Caching
LLM APIs that support prompt caching, ranked by cached-input price. Cache hits cost 50–99% less than fresh input — the biggest lever for cutting cost on repeated system prompts, RAG context, and long conversations.
MODELPRICEWATCH.COM · 2026-09-29
Cost calculator for this use casemonthly cost, top 3 models
- 1 GLM-4.6V-FlashX
- $—
- 2 Solar Mini 4
- $—
- 3 DeepSeek V4.1 Flash
- $—
Full ranking — top 20 models
list prices, USD per 1M tokens| Model | Cached input | Blended* | Input | Output | Context | Provider |
|---|---|---|---|---|---|---|
| 1GLM-4.6V-FlashX | $0.004 | $0.130 | $0.040 | $0.400 | 128K | Z.AI |
| 2Solar Mini 4 | $0.005 | $0.088 | $0.050 | $0.200 | 524K | Upstage |
| 3DeepSeek V4.1 Flash | $0.006 | $0.525 | $0.300 | $1.20 | 1M | DeepSeek |
| 4GPT-6 Luna | $0.010 | $0.200 | $0.100 | $0.500 | 1M | OpenAI |
| 5GLM-4.7-FlashX | $0.010 | $0.153 | $0.070 | $0.400 | 128K | Z.AI |
| 6GPT-OSS 120B | $0.015 | $0.262 | $0.150 | $0.600 | 128K | Fireworks |
| 7Solar Pro 3 | $0.015 | $0.262 | $0.150 | $0.600 | 131K | Upstage |
| 8Qwen3.8-Omni-Flash | $0.016 | $0.230 | $0.150 | $0.470 | 1M | Alibaba |
| 9Solar Pro 4 | $0.018 | $0.158 | $0.090 | $0.360 | 524K | Upstage |
| 10DeepSeek V4 Flash | $0.028 | $0.175 | $0.140 | $0.280 | 1M | Fireworks |
| 11GLM-4.5-Air | $0.030 | $0.425 | $0.200 | $1.10 | 128K | Z.AI |
| 12GLM-5.3-Flash | $0.030 | $0.237 | $0.150 | $0.500 | 1M | Z.AI |
| 13KAT-Coder-Air V2.5 | $0.030 | $0.259 | $0.148 | $0.593 | 256K | Kwaipilot |
| 14GPT-OSS 20B | $0.035 | $0.128 | $0.070 | $0.300 | 128K | Fireworks |
| 15Muse Glimmer 30B | $0.040 | $0.637 | $0.350 | $1.50 | 131K | Fireworks |
| 16NVIDIA Nemotron 3.5 Lightning | $0.040 | $0.110 | $0.080 | $0.200 | 256K | DeepInfra |
| 17DeepSeek V4 Pro | $0.044 | $1.98 | $1.32 | $3.96 | 1M | DeepSeek |
| 18GLM-4.6V | $0.050 | $0.450 | $0.300 | $0.900 | 128K | Z.AI |
| 19MiniMax-M2.7 | $0.060 | $0.525 | $0.300 | $1.20 | 205K | MiniMax |
| 20MiniMax-M3 | $0.060 | $0.525 | $0.300 | $1.20 | 1M | MiniMax |
Recent price movement in this ranking
price-only deltas · logged by the daily scan4 of the top 20 prompt caching models have re-priced since we began tracking · last checked .
| Model | Changes | Latest move | Provider |
|---|---|---|---|
| Solar Pro 4 | 1 | price rise on : $0.030/$0.120 → $0.090/$0.360 /Mtok | Upstage |
| GLM-5.3-Flash | 1 | price rise on : $0.075/$0.250 → $0.150/$0.500 /Mtok | Z.AI |
| DeepSeek V4 Pro | 1 | price rise on : $0.435/$0.870 → $1.32/$3.96 /Mtok | DeepSeek |
| MiniMax-M3 | 1 | price cut on : $0.600/$2.40 → $0.300/$1.20 /Mtok | MiniMax |
Most recently, Solar Pro 4 raised its price on — a sign pricing in this category is still moving, so re-check before committing to a long-term choice. See all recent price moves →
How models are selected
Models offering prompt caching, sorted by cached-input price per million tokens (cheapest cache reads first).
Prices are per million tokens (Mtok), sourced directly from — and linked to — official provider pricing pages. "Blended cost" is (3×input + 1×output) ÷ 4 — weighted toward input because real workloads read far more tokens than they generate.