ModelPriceWatch.com
Last scan 2026-08-15 Models tracked 205 Providers 31 Cheapest paid Granite 4.0 H Micro $0.017/Mtok in Every price links to its source

Best LLM APIs with Prompt Caching

LLM APIs that support prompt caching, ranked by cached-input price. Cache hits cost 50–99% less than fresh input — the biggest lever for cutting cost on repeated system prompts, RAG context, and long conversations.

36 models qualify top 20 shown sorted by cached-input price, cheapest first
1DeepSeek

DeepSeek V4 Flash

$0.003 /1M cached input

$0.175 blended · $0.140 in · $0.280 out

2DeepSeek

DeepSeek V4 Pro

$0.004 /1M cached input

$0.544 blended · $0.435 in · $0.870 out

3Upstage

Solar Pro 4

$0.006 /1M cached input

$0.052 blended · $0.030 in · $0.120 out

MODELPRICEWATCH.COM · 2026-08-15

Cost calculator for this use casemonthly cost, top 3 models

1 DeepSeek V4 Flash
$—
2 DeepSeek V4 Pro
$—
3 Solar Pro 4
$—

Full ranking — top 20 models

list prices, USD per 1M tokens
Top 20 models for Best LLM APIs with Prompt Caching, ranked by cached-input price, cheapest first
Model Cached input Blended* Input Output Context Provider
1DeepSeek V4 Flash $0.003 $0.175 $0.140 $0.280 1M DeepSeek
2DeepSeek V4 Pro $0.004 $0.544 $0.435 $0.870 1M DeepSeek
3Solar Pro 4 $0.006 $0.052 $0.030 $0.120 524K Upstage
4GLM-4.7-FlashX $0.010 $0.153 $0.070 $0.400 128K Z.AI
5GPT-OSS 120B $0.015 $0.262 $0.150 $0.600 128K Fireworks
6Solar Pro 3 $0.015 $0.262 $0.150 $0.600 131K Upstage
7GLM-4.5-Air $0.030 $0.425 $0.200 $1.10 128K Z.AI
8KAT-Coder-Air V2.5 $0.030 $0.259 $0.148 $0.593 256K Kwaipilot
9GPT-OSS 20B $0.035 $0.128 $0.070 $0.300 128K Fireworks
10GLM-4.6V $0.050 $0.450 $0.300 $0.900 128K Z.AI
11MiniMax-M2.7 $0.060 $0.525 $0.300 $1.20 205K MiniMax
12MiniMax-M3 $0.060 $0.525 $0.300 $1.20 1M MiniMax
13Inkling Small $0.060 $0.525 $0.300 $1.20 262K Thinking Machines
14Qwen3.7-Plus $0.080 $0.700 $0.400 $1.60 128K Fireworks
15Kimi K2.5 $0.100 $1.20 $0.600 $3.00 262K Moonshot
16GLM-4.5 $0.110 $1.00 $0.600 $2.20 128K Z.AI
17GLM-4.6 $0.110 $1.00 $0.600 $2.20 200K Z.AI
18GLM-4.7 $0.110 $1.00 $0.600 $2.20 128K Z.AI
19NVIDIA Nemotron 3 Ultra $0.120 $1.05 $0.600 $2.40 128K Fireworks
20Qwen3.7-Max $0.130 $1.88 $1.25 $3.75 1M Together
Ranking basis: cached-input price, cheapest first. * Blended = (3×input + 1×output) ÷ 4. Every price links to its source on the model page.

Recent price movement in this ranking

price-only deltas · logged by the daily scan

1 of the top 20 prompt caching models has re-priced since we began tracking · last checked Aug 15, 2026.

Models in this Best LLM APIs with Prompt Caching ranking that have re-priced since tracking began, most recent move first
Model Changes Latest move Provider
MiniMax-M3 1 price cut on Jun 17, 2026: $0.6/$2.4 → $0.3/$1.2 /Mtok MiniMax
Changes = distinct price moves logged since tracking began; in/out prices are $ per 1M tokens. MODELPRICEWATCH.COM · 2026-08-15

Most recently, MiniMax-M3 cut its price on Jun 17, 2026 — a sign pricing in this category is still moving, so re-check before committing to a long-term choice. See all recent price moves →

How models are selected

Models offering prompt caching, sorted by cached-input price per million tokens (cheapest cache reads first).

Prices are per million tokens (Mtok), sourced directly from — and linked to — official provider pricing pages. "Blended cost" is (3×input + 1×output) ÷ 4 — weighted toward input because real workloads read far more tokens than they generate.

Other use case rankings