The LLM Price Index — Methodology
A single, transparent number for the cost of frontier intelligence — and exactly how it's built.
The LLM Price Index · as of Aug 15, 2026
$4.32/1M tok
per million tokens · blended 3:1 input:output · equal-weight across 10 frontier flagships · -5.5% since Feb 23, 2026
The level is computed fresh on every build from verified prices. The historical trend fills in as price history accrues — see History & honesty.
What it measures
The LLM Price Index is the equal-weighted average blended cost per million tokens across a fixed basket of frontier-flagship language models — one current, general-purpose flagship from each major lab. It answers one question in a single number: what does it cost to run a million tokens through today's best models?
It exists so the cost of frontier AI can be quoted — the way the Consumer Price Index or a stock index lets a whole market be referenced with one figure. A 156-row table can't travel into an article, a tweet, or an AI answer; one named number can.
The basket
Constituents are fixed and explicit so the index is stable and auditable — it changes only when we deliberately rebalance (see below), never silently. Today's basket is 10 constituents — one current, general-purpose flagship per lab we track:
| Model | Provider | Input $/1M | Output $/1M | Blended $/1M | Weights |
|---|---|---|---|---|---|
| GPT-5.6 Sol | OpenAI | $5.00 | $30.00 | $11.25 | Closed |
| Claude Opus 5 | Anthropic | $5.00 | $25.00 | $10.00 | Closed |
| Gemini 3.1 Pro | $2.00 | $12.00 | $4.50 | Closed | |
| Grok 4.6 | xAI | $2.00 | $6.00 | $3.00 | Closed |
| DeepSeek V4 Pro | DeepSeek | $0.435 | $0.870 | $0.544 | Closed |
| Qwen3.8-Max | Alibaba | $2.00 | $6.00 | $3.00 | Closed |
| Mistral Large 3 | Mistral | $0.500 | $1.50 | $0.750 | Closed |
| GLM-5.2 | Z.AI | $1.40 | $4.40 | $2.15 | Open |
| Kimi K3 | Moonshot | $3.00 | $15.00 | $6.00 | Open |
| Muse Spark 1.2 | Meta | $1.25 | $4.25 | $2.00 | Closed |
Weighting & the 3:1 blend
Equal-weight. Every constituent counts the same. This is the most transparent and least gameable choice — we don't have reliable, neutral usage data to weight by popularity, and inventing weights would invite bias. Equal-weight means the index reflects the price landscape, not our guess at market share.
Blended 3:1 input:output. Output tokens cost more than input tokens, so a single "price" needs an assumed mix. A typical request reads far more than it writes (a prompt plus context, then a shorter answer), so we weight input 3× output:
blended = (3 × input + 1 × output) ÷ 4
This assumption is held constant across every model and every date, so it never distorts comparisons or the trend.
Sub-indices
The headline is one number; these cut it three ways that each tell a different story:
- Cheapest frontier token — the lowest blended cost in the basket: $0.544/Mtok (DeepSeek V4 Pro). How little frontier-class capability can cost.
- Cost per intelligence point — blended cost ÷ average benchmark score, averaged across the basket: $0.062 per point (coverage 7/10); best value: DeepSeek V4 Pro. Price adjusted for quality.
- Open vs closed spread — closed-weights flagships average $4.38/Mtok vs open-weights $4.08/Mtok — a spread of $0.310. The premium for proprietary models.
The Intelligence Cost Curve
The flagship index above is a ceiling — the price of the frontier, which is sticky by construction. Beside it we plot a floor: the cheapest model that still clears a fixed capability bar, over time. As cheaper capable models launch, the floor falls; the gap between the two lines is the real deflation story.
Floor constructionmetric · bar · provenance · dedup · blend
Today's floor
$0.113/1M tok
Today the floor sits at DeepSeek V4 Flash — about 38.2× below the $4.32 flagship ceiling.
The floor is plotted only from Mar 1, 2026 — the first date on which a qualifying model had both a verified score and a listed price. We do not synthesize the floor before that. Benchmark scores are snapshotted on every refresh into an append-only history (idempotent by content); where a score's first-known date predates that pipeline, it is reconstructed from the recorded source date and labelled as such — never presented as a fresh capture.
What a model must clear
- Metric — GPQA Diamond (raw %). A single, absolute 0–100 score with a fixed meaning, not a percentile rank against the current population. This is deliberate: a percentile composite has artifacts (a small model can rank-inflate above a flagship), and its meaning drifts every time a model is added. A raw GPQA score cannot.
- Bar — GPQA Diamond ≥ 70 ("GPT-4-class general reasoning"). Anchored to a named, frozen reference — GPT-4.1 (GPQA Diamond 70.0) — so the question stays concrete ("what does GPT‑4‑class intelligence cost today?") and the bar never drifts once locked.
- Provenance — external only (self_reported excluded). Only externally-sourced scores count; self-reported and vendor-card-only numbers are excluded, and any score outside the metric's valid 0–100 range is dropped before it can define the frontier max. No score may exceed the best externally-verified model (96.2).
- Dedup — base model (min blended across hosted SKUs). A model hosted at many providers/tiers collapses to its cheapest live SKU, so a base model is represented once at its lowest verified price.
- Blend — 3:1 input:output, the same convention as the ceiling, so the two lines are apples-to-apples.
The provenance rule, workedvendor-reported scores do not set the floor
A worked example of the provenance rule. Moonshot's Kimi K3 is a basket constituent above, and the only GPQA Diamond figure on record for it — 93.5 — comes from Moonshot's own launch materials. It is therefore excluded from the curve: a self-reported score that would sit near the top of the verified range is exactly the kind of number that could define a false floor, and no vendor grades its own homework here. This is a rule about the source, not the model — if an independent evaluator (Artificial Analysis, Vellum) publishes a GPQA Diamond score for K3, it becomes eligible automatically, with no change to the code or the bar. We would rather show a conservative floor than a flattering one.
History & honesty
The trend is a chain-linked index. Because frontier flagships launch at different times, at any past date only part of the basket has a listed price — so each step measures the price move only across the constituents present in both adjacent periods (the matched sample) and compounds those moves. A new flagship joining the basket therefore never creates a phantom jump. The chained level is rescaled so today's point equals the headline above. The index has moved -5.5% between Feb 23, 2026 and Aug 15, 2026.
Every price in the index is a fact sourced from an official provider page and re-verified on our normal refresh cycle. The index is recomputed on every build, so the level is never stale.
Rebalancing
The basket changes only by deliberate decision — when a lab ships a new flagship that supersedes its current constituent, when a constituent is retired, or when a lab we did not previously cover ships a frontier flagship and earns a slot. Rebalances are logged here and marked on the trend chart. We do not add or drop models to move the number.
There are two distinct kinds of change, and we label them differently because they mean different things:
- Succession — a lab's slot moves to that lab's newer flagship. The price step from the retired model to its successor is a real move in that lab's flagship cost, so it chains through the trend. A like-for-like succession at the same list price leaves the level unchanged.
- Composition — a lab with no slot gets one. The chain-linked trend is unaffected by design (a newly-added constituent is not in the matched sample on the step it appears, so it cannot create a phantom jump), but the headline is an equal-weight mean over a larger basket, so the published number moves. That is a change in what is measured, not a market move, and we say so rather than letting it read as a price event.
On 25 July 2026 we opened a ninth slot for Moonshot (Kimi K3, $3.00/$15.00 per 1M, 1M-token context, open weights). Moonshot had no slot at all, so its flagship could never enter the index however current it was — a coverage gap, not a price signal. Adding it moved the headline from $4.07 to $4.25 as a composition change.
Also on 25 July 2026, Alibaba's slot moved by succession from Qwen3-Max to its true current flagship, Qwen3.7-Max ($2.50/$7.50 per 1M list, 1M-token context) — a real +56% step in that lab's flagship blended cost that had gone unrecorded for two months, lifting the headline from $4.25 to $4.39. It is chained into the trend as a market move.
Maker vs host. That two-month lag is worth explaining, because it is the kind of blind spot we now guard against. Qwen3.7-Max is Alibaba's flagship, but for a while it appeared in our dataset only as a listing under a reseller (Together) — so its record carried the host's name, not the maker's. A basket that tracks one flagship per lab keys on the maker, so a successor filed under a host can slip between "is this slot superseded?" (which compared makers) and "does an uncovered lab ship a flagship?" (which skips resellers). Our freshness gate now resolves the maker, not the provider, on both questions, so a maker's newer flagship is caught even when it is listed only under a host.
How a missing lab surfaces. Our basket-freshness gate used to ask only "is this slot's occupant superseded?", which is structurally blind to a lab holding no slot. It now also asks the inverted question: does any lab without a slot ship a current, priced, general-purpose flagship? If so it names the candidate at build time. Resellers relisting someone else's open-weights model, stale-vintage flagships, and specialised (search, embedding, image, coding) products are excluded. The gate surfaces candidates; a human decides — we will not let an automated rule silently change what the index measures.
On 1 August 2026 that inverted question fired for the first time: we opened a tenth slot for Meta. Muse Spark 1.1 (released 2026-07-09; $1.25/$4.25 per 1M first-party, 1M-token context) is Meta's first proprietary paid-API model and its current general-purpose flagship, but the lab held no slot, so however current the model was it could never enter. This is a composition change, not a market move: the chain-linked trend is unaffected by design, and the addition moved the headline from $4.66 to $4.39. Because the headline is an equal-weight mean and the whole chained history is rescaled onto the new ten-lab basis, every previously published level is restated too — which is exactly why the next section exists.
On 3 August 2026 Alibaba made Qwen3.8-Max ($2.00/$6.00 per 1M list, one flat tier across the full 1M-token context) generally available, superseding Qwen3.7-Max in that lab's slot. The handover is dated 4 August 2026 — the first day we publish Qwen3.8-Max as the Alibaba occupant — not the 3 August GA date, because 3 August had already been published and archived with Qwen3.7-Max in the slot, and we do not restate a level we have already published. From the handover the index carries a real −20% step in that lab's flagship blended cost, $3.75 down to $3.00, chained into the trend as a market move. Both ends of that step are list prices, the basis this index publishes: the retired Qwen3.7-Max was also running a limited-time 50% promotion ($1.25/$3.75 effective), which the index does not track. Because this is a succession rather than a composition change, it lands as a dated step and leaves the published level history where it stands — the basket stays at ten labs, and the dated levels you may have cited remain frozen in the append-only archive.
What a level means — and what we promise
The Price Index is a trend instrument published on the current basis, the same convention consumer-price and equity indices use after a basket revision. Concretely, that means three promises:
- The trend is the invariant. The percent change we publish is chain-linked over matched samples, so it is unaffected by composition changes. When we say the index has moved -5.5% since its first reading, that statement survives every rebalance.
- Levels always speak today's basis. When a lab earns a new slot, the entire published history is restated onto the new basis — on the current ten-lab basis the index opened its record at $4.57/Mtok, where the superseded nine-lab basis had opened at $4.84. A past level you read here today is the current-basis restatement, not the number that was on this page at the time. We mark every such change on the chart and log it above; we never let a composition change read as a price event.
- Dated citations stay reproducible. Each deploy archives the headline exactly as published that day, on that day's basis, to an append-only record that is never restated: /api/v1/price-index-levels.json. If you cited a level on a given date, that row is your receipt — even after a later rebalance re-levels the live history. Our dated reports are likewise frozen at their stated basis and corrected only with visible notes.
Rule of thumb: quote trends from the live index; quote levels with a date, and check dated levels against the as-published archive. Comparing levels across two bases is the one misuse the design cannot prevent, so we ask you not to — and give you the archive so you never have to.
Use the data
Every constituent has its own model page with verified pricing and the official source link. The full dataset is free and CORS-open at the API. Cite the LLM Price Index freely — attribution to modelpricewatch.com is appreciated.