ModelPriceWatch.com
Last scan 2026-09-29 Models tracked 272 Providers 35 Cheapest paid Granite 4.0 H Micro $0.017/Mtok in Every price links to its source

Best LLM APIs with Prompt Caching

LLM APIs that support prompt caching, ranked by cached-input price. Cache hits cost 50–99% less than fresh input — the biggest lever for cutting cost on repeated system prompts, RAG context, and long conversations.

56 models qualify top 20 shown sorted by cached-input price, cheapest first
1Z.AI

GLM-4.6V-FlashX

$0.004 /1M cached input

$0.130 blended · $0.040 in · $0.400 out

2Upstage

Solar Mini 4

$0.005 /1M cached input

$0.088 blended · $0.050 in · $0.200 out

3DeepSeek

DeepSeek V4.1 Flash

$0.006 /1M cached input

$0.525 blended · $0.300 in · $1.20 out

MODELPRICEWATCH.COM · 2026-09-29

Cost calculator for this use casemonthly cost, top 3 models

1 GLM-4.6V-FlashX
$—
2 Solar Mini 4
$—
3 DeepSeek V4.1 Flash
$—

Full ranking — top 20 models

list prices, USD per 1M tokens
Top 20 models for Best LLM APIs with Prompt Caching, ranked by cached-input price, cheapest first
Model Cached input Blended* Input Output Context Provider
1GLM-4.6V-FlashX $0.004 $0.130 $0.040 $0.400 128K Z.AI
2Solar Mini 4 $0.005 $0.088 $0.050 $0.200 524K Upstage
3DeepSeek V4.1 Flash $0.006 $0.525 $0.300 $1.20 1M DeepSeek
4GPT-6 Luna $0.010 $0.200 $0.100 $0.500 1M OpenAI
5GLM-4.7-FlashX $0.010 $0.153 $0.070 $0.400 128K Z.AI
6GPT-OSS 120B $0.015 $0.262 $0.150 $0.600 128K Fireworks
7Solar Pro 3 $0.015 $0.262 $0.150 $0.600 131K Upstage
8Qwen3.8-Omni-Flash $0.016 $0.230 $0.150 $0.470 1M Alibaba
9Solar Pro 4 $0.018 $0.158 $0.090 $0.360 524K Upstage
10DeepSeek V4 Flash $0.028 $0.175 $0.140 $0.280 1M Fireworks
11GLM-4.5-Air $0.030 $0.425 $0.200 $1.10 128K Z.AI
12GLM-5.3-Flash $0.030 $0.237 $0.150 $0.500 1M Z.AI
13KAT-Coder-Air V2.5 $0.030 $0.259 $0.148 $0.593 256K Kwaipilot
14GPT-OSS 20B $0.035 $0.128 $0.070 $0.300 128K Fireworks
15Muse Glimmer 30B $0.040 $0.637 $0.350 $1.50 131K Fireworks
16NVIDIA Nemotron 3.5 Lightning $0.040 $0.110 $0.080 $0.200 256K DeepInfra
17DeepSeek V4 Pro $0.044 $1.98 $1.32 $3.96 1M DeepSeek
18GLM-4.6V $0.050 $0.450 $0.300 $0.900 128K Z.AI
19MiniMax-M2.7 $0.060 $0.525 $0.300 $1.20 205K MiniMax
20MiniMax-M3 $0.060 $0.525 $0.300 $1.20 1M MiniMax
Ranking basis: cached-input price, cheapest first. * Blended = (3×input + 1×output) ÷ 4. Every price links to its source on the model page.

Recent price movement in this ranking

price-only deltas · logged by the daily scan

4 of the top 20 prompt caching models have re-priced since we began tracking · last checked .

Models in this Best LLM APIs with Prompt Caching ranking that have re-priced since tracking began, most recent move first
Model Changes Latest move Provider
Solar Pro 4 1 price rise on : $0.030/$0.120 → $0.090/$0.360 /Mtok Upstage
GLM-5.3-Flash 1 price rise on : $0.075/$0.250 → $0.150/$0.500 /Mtok Z.AI
DeepSeek V4 Pro 1 price rise on : $0.435/$0.870 → $1.32/$3.96 /Mtok DeepSeek
MiniMax-M3 1 price cut on : $0.600/$2.40 → $0.300/$1.20 /Mtok MiniMax
Changes = distinct price moves logged since tracking began; in/out prices are $ per 1M tokens. MODELPRICEWATCH.COM · 2026-09-29

Most recently, Solar Pro 4 raised its price on — a sign pricing in this category is still moving, so re-check before committing to a long-term choice. See all recent price moves →

How models are selected

Models offering prompt caching, sorted by cached-input price per million tokens (cheapest cache reads first).

Prices are per million tokens (Mtok), sourced directly from — and linked to — official provider pricing pages. "Blended cost" is (3×input + 1×output) ÷ 4 — weighted toward input because real workloads read far more tokens than they generate.

Other use case rankings