Best LLM APIs for Long Context Windows
LLM APIs with the largest context windows. Compare models that support 100K+ tokens for document analysis, codebase processing, and long conversations.
MODELPRICEWATCH.COM · 2026-08-09
Cost calculator for this use casemonthly cost, top 3 models
- 1 Llama 4 Scout
- $—
- 2 Gemini 3.1 Pro
- $—
- 3 GPT-5.6 Luna
- $—
Full ranking — top 20 models
list prices, USD per 1M tokens| Model | Blended* | Input | Output | Context | Provider |
|---|---|---|---|---|---|
| 1Llama 4 Scout | $0.168 | $0.110 | $0.340 | 10M | Meta |
| 2Gemini 3.1 Pro | $4.50 | $2.00 | $12.00 | 2M | |
| 3GPT-5.6 Luna | $0.450 | $0.200 | $1.20 | 1M | OpenAI |
| 4GPT-5.6 Terra | $4.50 | $2.00 | $12.00 | 1M | OpenAI |
| 5GPT-5.5 | $11.25 | $5.00 | $30.00 | 1M | OpenAI |
| 6GPT-5.6 Sol | $11.25 | $5.00 | $30.00 | 1M | OpenAI |
| 7GPT-5.4 Pro | $67.50 | $30.00 | $180.00 | 1M | OpenAI |
| 8GPT-5.5 Pro | $67.50 | $30.00 | $180.00 | 1M | OpenAI |
| 9Gemini 3 Flash Preview | $1.13 | $0.500 | $3.00 | 1M | |
| 10Muse Spark 1.1 | $2.00 | $1.25 | $4.25 | 1M | Meta |
| 11Muse Spark 1.2 | $2.00 | $1.25 | $4.25 | 1M | Meta |
| 12Kimi K3 | $6.00 | $3.00 | $15.00 | 1M | Moonshot |
| 13Qwen3.7-Flash | $0.055 | $0.030 | $0.130 | 1M | Alibaba |
| 14Qwen-Turbo | $0.088 | $0.050 | $0.200 | 1M | Alibaba |
| 15Qwen-Flash | $0.138 | $0.050 | $0.400 | 1M | Alibaba |
| 16DeepSeek V4 Flash | $0.175 | $0.140 | $0.280 | 1M | DeepSeek |
| 17Llama 4 Maverick | $0.350 | $0.200 | $0.800 | 1M | DeepInfra |
| 18MiniMax-M3 | $0.525 | $0.300 | $1.20 | 1M | MiniMax |
| 19DeepSeek V4 Pro | $0.544 | $0.435 | $0.870 | 1M | DeepSeek |
| 20Qwen3.6-Flash | $0.563 | $0.250 | $1.50 | 1M | Alibaba |
Recent price movement in this ranking
price-only deltas · logged by the daily scan3 of the top 20 long context models have re-priced since we began tracking · last checked Aug 9, 2026.
| Model | Changes | Latest move | Provider |
|---|---|---|---|
| GPT-5.6 Luna | 1 | price cut on Aug 1, 2026: $1/$6 → $0.2/$1.2 /Mtok | OpenAI |
| GPT-5.6 Terra | 1 | price cut on Aug 1, 2026: $2.5/$15 → $2/$12 /Mtok | OpenAI |
| MiniMax-M3 | 1 | price cut on Jun 17, 2026: $0.6/$2.4 → $0.3/$1.2 /Mtok | MiniMax |
Most recently, GPT-5.6 Luna cut its price on Aug 1, 2026 — a sign pricing in this category is still moving, so re-check before committing to a long-term choice. See all recent price moves →
How models are selected
Models with 100K+ token context windows, sorted by context size (largest first); models tied on context size list the maker's own row first, then cheapest blended price.
Prices are per million tokens (Mtok), sourced directly from — and linked to — official provider pricing pages. "Blended cost" is (3×input + 1×output) ÷ 4 — weighted toward input because real workloads read far more tokens than they generate.