ModelPriceWatch.com
Last scan 2026-09-05 Models tracked 248 Providers 32 Cheapest paid Granite 4.0 H Micro $0.017/Mtok in Every price links to its source

Mercury 2

by Inception

Current budget cheap tier

Today's price · per 1M tokens

Input

$0.250

per 1M tokens

Output

$0.750

per 1M tokens

Blended

$0.375

blended $/1M — 3:1 weighted input:output

Cached input

$0.025

10% of input — prompt caching

How this price is scoped: Inception publishes this rate per TOKEN, not per 1M: its own API root (api.inceptionlabs.ai/v1/models) serves prompt 0.00000025 and completion 0.00000075, which is $0.25 / $0.75 per 1M tokens, with cache reads at 0.000000025 ($0.025 per 1M) and cache writes free. Read unauthenticated on 2026-09-01 and again on 2026-09-02, byte-identical both times (973 bytes, md5 a835588b3e712d66d80ab0166a1b8fba). The vendor's own record carries "openrouter": {"slug": "inception/mercury-2"}, so the join to the gateway id is Inception's claim rather than ours. Inception's HTML pricing page is behind a standing bot challenge and is not our source.

Source: official Inception pricing · read MODELPRICEWATCH.COM · 2026-09-05

Price receipt

We read Inception's own model API on and found mercury-2 listed with a price in its own record — that read is where the number above comes from, recorded as a new model. Inception states this rate per token; the per-1M figures above are that rate multiplied by 1,000,000 — no conversion or estimate is involved. We keep the page text we read; its content hash is 998c31a236.

Overview

The first diffusion large language model (dLLM) sold as a commercial API. Mercury 2 generates tokens by iterative denoising rather than left-to-right decoding, which Inception reports as 5-10x faster than speed-optimised autoregressive models at comparable quality. 128K context, 50K max output, tools and structured outputs. $0.25 / $0.75 per 1M tokens.

Capabilities

struck through = not supported
Input 1/5
Text Image Audio Video PDF
Output 1/5
Text Image Audio Video Embedding
Features 3/9
Prompt caching Reasoning Coding Fast inference Long context Open weights Multimodal Web search Realtime

Benchmark performance

accuracy % · higher is better
Percentile vs all tracked models 12.4th2 independent measurements
Percentile per $/Mtok 32.6

We hold no per-benchmark accuracy scores for this model yet, so it has no accuracy average. That is a gap in our coverage, not a sign the model is untested — independent evaluators often publish a composite index for a new model long before releasing its per-benchmark numbers. The percentile above is its standing across the independent composites below.

Independent composite scores
  • AA Agentic Index: 4.1 (Tool use, planning, autonomy)
  • AA-Omniscience Index: -50.7 (−100–100)
How it stacks up
  • Ranks #153 of 180 comparably-measured models by percentile score, across 2 independent measurements
  • Ranks #92 of 180 comparably-tested models by normalized performance per dollar

Source: Artificial Analysis · updated · See full rankings →

Specifications

Provider
Inception
Context window
128K tokens
Modality
text
Parameters
Proprietary
Open source
No — proprietary
Released
Status
Current
Last updated
Tags
budgetfast

Availability verified: listed on Inception's own page