State of LLM Pricing — September 2026
What a million tokens costs, from verified provider list prices. All figures as of September 1, 2026; this edition does not change after publication — the live index does. Previous edition: August 2026.
Key findings
data through 2026-09-01- Frontier intelligence cost $4.14 per million tokens on September 1, 2026. The LLM Price Index — the equal-weight blended price across 10 current flagship models, one per major lab — closed the month at $4.14/Mtok, down 5.7% from $4.39 on August 1 and down 9.4% from $4.57 at the first reading on February 23, 2026.
- The sticky ceiling cracked. The August edition reported that across 40 readings no basket constituent had ever changed its own list price. In August two did: DeepSeek tripled the peak rate on V4 Pro on August 16 (+264% blended), and OpenAI cut GPT-5.6 Sol by 29% on August 21 under a promotional label. With the Qwen3.8-Max succession on August 4, the level moved on three dates in one month — more than in the preceding five months combined.
- The floor held only because a reseller held it. DeepSeek raised its own V4 Flash peak rate from $0.14/$0.28 to $0.44/$1.32 on August 16 (blended $0.175 → $0.66), but DeepInfra kept serving the same model at $0.09/$0.18, so the cheapest model clearing a fixed GPT-4-class bar (GPQA Diamond ≥ 70) stayed at $0.113/Mtok. GPT-4-class capability was 37x cheaper than the flagship ceiling on September 1.
- "Frontier flagship" spans a 13x price range, down from 21x. Both ends moved inward: the top fell from $11.25/Mtok (GPT-5.6 Sol) to $10.00 (Claude Opus 5) as Sol took its cut, and the bottom rose from $0.544 (DeepSeek V4 Pro) to $0.75 (Mistral Large 3) as DeepSeek moved to peak pricing.
- Price per unit of measured intelligence averaged $0.0600 on September 1 (blended $/Mtok ÷ benchmark score, over the 7 of 10 constituents with a published score), down 6% from $0.0637 in August. The best value changed hands: DeepSeek V4 Pro's reprice took it from $0.0076 to $0.0269 per point, and Mistral Large 3 at $0.0140 per point became the cheapest scored flagship — 4.3x better than the frontier average.
The index level
chain-linked · composition-neutralThe index entered August at $4.39/Mtok on the ten-lab basket and left it at $4.14. It moved on three dates:
| Date | Event | What changed | Level $/1M |
|---|---|---|---|
| Aug 1, 2026 | — | Opening level, ten-lab basket | $4.39 |
| Aug 4, 2026 | Rebalance | Alibaba's slot: Qwen3.7-Max ($3.75 blended) → Qwen3.8-Max ($3.00, list $2/$6 flat across the full 1M context) | $4.32 |
| Aug 16, 2026 | Reprice ↑ | DeepSeek V4 Pro: flat $0.435/$0.87 → peak $1.32/$3.96, off-peak half ($0.544 → $1.98 blended, +264%) | $4.46 |
| Aug 21, 2026 | Reprice ↓ | GPT-5.6 Sol: $5/$30 → $4/$20, labelled promotional ($11.25 → $8.00 blended, −29%) | $4.14 |
| Sep 1, 2026 | — | Closing level, unchanged since August 21 | $4.14 |
On August 4 the level stepped down to $4.32 when Alibaba's slot passed from Qwen3.7-Max to the cheaper Qwen3.8-Max — a succession that changes the slot's price, so it chains into the trend as a real move. On August 16 it stepped up to $4.46 when DeepSeek replaced V4 Pro's flat rate with a peak/off-peak schedule: $1.32 input / $3.96 output per million during the seven peak hours (01:00–04:00 and 06:00–10:00 UTC), exactly half outside them. The index tracks the peak rate as the list price, because DeepSeek defines off-peak as a discount from peak and a caller who does not schedule around the clock needs a ceiling, not a floor; even the off-peak rate ($0.99 blended) is 82% above the old flat price. On August 21 it stepped down to $4.14 when OpenAI cut GPT-5.6 Sol from $5/$30 to $4/$20 — the first list-price cut by any constituent in tracked history. OpenAI labels the new rate promotional, "available at least through November 21, 2026", and publishes no date on which it reverts; the index takes list prices as printed, so the cut is in.
A note on the basis. Every level above is stated on the ten-lab basket, which did not change composition in August, so nothing was rescaled this month: the August 1 level of $4.39 is the same number the August edition published, and the February-to-September trend (−9.4%) is read on one basis end to end. The index is chain-linked; a new lab entering (as Meta did on August 1) restates the level history without changing the measured trend, but no lab entered in August.
Two of August's three moves were reprices. That changes the finding, not the method. The index still moves rarely — five dates in 71 readings since February — but "frontier list prices never move" is no longer a description of the record. Both reprices carry conditions, a time-of-day schedule and a promotional floor date, so in two of ten slots the published list price is now the top of a range rather than a single number.
The basket, September 1
one current flagship per major lab| Model | Provider | Blended* $/1M | Weights |
|---|---|---|---|
| Claude Opus 5 | Anthropic | $10.00 | Closed |
| GPT-5.6 Sol † | OpenAI | $8.00 | Closed |
| Kimi K3 | Moonshot | $6.00 | Open |
| Gemini 3.1 Pro | $4.50 | Closed | |
| Qwen3.8-Max | Alibaba | $3.00 | Closed |
| Grok 4.6 | xAI | $3.00 | Closed |
| GLM-5.3 ‡ | Z.AI | $2.15 | Closed |
| Muse Spark 1.2 | Meta | $2.00 | Closed |
| DeepSeek V4 Pro § | DeepSeek | $1.98 | Closed |
| Mistral Large 3 | Mistral | $0.75 | Closed |
The spread from the priciest constituent to the cheapest narrowed from 21x to 13x, and it narrowed from both ends: Sol's cut lowered the top of the range from $11.25 to Claude Opus 5's $10.00, and DeepSeek's move to peak pricing lifted the bottom from $0.544 to Mistral Large 3's $0.75. With GLM-5.3 arriving closed, Kimi K3 is the basket's only open-weights constituent, so the August edition's open-versus-closed comparison (a 10% proprietary premium) has no counterpart this month: the nine closed-weights flagships averaged $3.93/Mtok against $6.00 for the one open-weights flagship, a comparison of nine numbers to one.
What GPT-4-class capability costs
cheapest model clearing GPQA Diamond ≥ 70Hold the capability bar fixed and August's headline is what did not happen. The cheapest model with an externally measured GPQA Diamond score of 70 or better — roughly GPT-4-class general reasoning — stayed at $0.113/Mtok for the whole month, its fifth step since March still standing:
| Floor taken by | Date | Blended $/1M |
|---|---|---|
| o4-mini | Mar 1, 2026 | $1.93 |
| Grok 4.3 | May 15, 2026 | $1.56 |
| DeepSeek V4 Pro | Jun 23, 2026 | $0.544 |
| GPT OSS 120B | Jun 29, 2026 | $0.262 |
| DeepSeek V4 Flash (via DeepInfra) | Jul 25, 2026 | $0.113 |
Underneath the flat line, the floor nearly broke upward. On August 16 DeepSeek raised the first-party peak rate for V4 Flash — the model that sets the floor — from $0.14/$0.28 to $0.44/$1.32, a blended $0.175 → $0.66 (3.8x). The floor did not move because the curve reads the cheapest listed price for a base model across every host that serves it, and DeepInfra kept V4 Flash at $0.09/$0.18 — $0.113 blended, 5.9x below the maker's own peak rate. That is a different kind of floor from the four before it: set not by the lab that trained the model but by a third party that hosts it, and therefore a price one vendor can withdraw. No new model undercut it in August; the month's two sub-$0.25 launches, GLM-5.3-Flash ($0.119 blended, promotional) and Qwen3.8-Flash ($0.23), have no externally measured GPQA Diamond score yet and neither is cheaper than $0.113.
After the cutoff
September 1–4, 2026 · stated so this page and the live index agreeThis edition is fixed at September 1. Two events in the following 72 hours are large enough that a reader arriving from the live index needs them named here; their full accounting belongs to the October edition.
- September 1: Anthropic launched Claude Fable 5.1 (and its limited-availability counterpart Mythos 5.1) at $10/$50 with cached input at $0.25 — 2.5% of the input rate, the deepest cache discount on any flagship rate card tracked. Anthropic's basket slot remains Claude Opus 5 ($5/$25), the lab's general-purpose flagship.
- September 3: OpenAI released GPT-6 Astra at $10/$50 ($20.00 blended) and named it its default flagship. It took the OpenAI slot from GPT-5.6 Sol on September 4, the first day the successor was published in the index, and the level stepped up to $5.34 — +29.0% in a day, the largest single move in the index's history, and the first time the index has stood above its February 23 first reading (+16.8%). The frontier spread reopened to 27x ($20.00 to $0.75). At a published score of 96, Astra costs $0.208 per point of measured intelligence, 3.5x the September 1 frontier average.
Two $10/$50 flagships in 72 hours, after a month in which the ceiling fell. Whether the ceiling has turned is a question for the October edition; the September 1 figures above stand as published.
Cite today's level
"Frontier intelligence costs $5.34 per million tokens — 16.8% since Feb 23, 2026."
The LLM Price Index — ModelPriceWatch, 2026-09-04 · modelpricewatch.com/price-index/
The data behind this report
open, attributable, re-verifiableEvery figure above is computed from the ModelPriceWatch tracker: 244 model rows (199 currently sold) across 31 providers, 12,807 price snapshots spanning February 2024 to September 1, 2026, each price traced to the provider's official pricing page. The full dataset is public:
Hugging Face dataset →
The complete token price history plus the current models table and the index series, CC-BY-4.0.
Price Index API →
The live index level, constituents, series, and the dated citation string. Free, CORS-open.
Methodology →
Basket rules, the 3:1 in:out blend, chain-linking, and rebalance policy.
Method, in one paragraph. The LLM Price Index is the equal-weighted average blended price ((3×input + output) ÷ 4 per million tokens) across a fixed basket of current general-purpose flagships, one per major lab. The basket changes only on a deliberate, dated rebalance; the trend is chain-linked from matched samples, so adding a lab never fabricates a price move. Where a vendor publishes a time-of-day schedule the peak rate is the list price; where a vendor labels a rate promotional it is taken as printed and flagged. The capability floor uses one absolute metric (GPQA Diamond, raw %) with recorded provenance per score, read across every host that serves a base model. Prices are list prices from official provider pages, re-verified on a rolling schedule — never estimated, never imputed. This report states only what that data shows. This edition is fixed at September 1, 2026 — for current numbers use the live index.