Inception
Commercial provider · 2 models tracked · Founded 2024
Palo Alto lab building diffusion large language models (dLLMs), founded in 2024 by the Stanford, UCLA and Cornell researchers behind the diffusion-model line of work. Its Mercury family generates and refines many tokens in parallel by iterative denoising rather than decoding left to right, which Inception sells on latency. Mercury 2 is sold first-party at $0.25/1M input and $0.75/1M output; Mercury 2.5 lists at $0.20/1M input and $0.75/1M output and is currently on an undated 80%-off promotion.
Inception pricing at a glanceSeptember 2026 · $ per 1M tokens
Inception API pricing (September 2026): 2 current models range from $0.200 to $0.250 per 1M input tokens and $0.750 to $0.750 per 1M output tokens. The cheapest paid model is Mercury 2.5 at $0.200/1M input; the priciest is Mercury 2 at $0.250/1M input / $0.750 output. Every price links to Inception's official pricing page and refreshes twice daily.
Today's Inception prices
2 models · cheapest blended firstSorted by blended cost (cheapest first). Prices per 1M tokens, September 2026 — every price links to Inception's official pricing page.
| Model | Blended* | Input | Output | Cached in | Relative cost | Context | Status |
|---|---|---|---|---|---|---|---|
| Mercury 2.5 | $0.338 |
$0.200 | $0.750 | $0.020 | 260K | Current | |
| Mercury 2 | $0.375 |
$0.250 | $0.750 | $0.025 | 128K | Current |
Quick stats
- Models tracked
- 2
- Type
- commercial
- Founded
- 2024
- Cheapest model
- Mercury 2.5
- Cheapest blended
- $0.338/M