ModelPriceWatch.com
Last scan 2026-09-10 Models tracked 249 Providers 32 Cheapest paid Granite 4.0 H Micro $0.017/Mtok in Every price links to its source

Mercury 2.5 vs Gemini 3.5 Flash-Lite

Side-by-side comparison of API pricing, specs, benchmarks, and capabilities

Mercury 2.5 is 60.3% cheaper on blended cost ($0.338 vs $0.850/Mtok)
Specification
Mercury 2.5Intro price
Overview
StatusCurrent budget Current budget
Released
Pricing per million tokens
Input $0.200/Mtok $0.300/Mtok
Output $0.750/Mtok $2.50/Mtok
Blended avg $0.338/Mtok $0.850/Mtok
Cached input $0.020/Mtok $0.030/Mtok
Price basis These are Inception's LIST prices, per 1M tokens. Mercury 2.5 is currently on an 80%-off promotion, so the rate you actually pay today is $0.04 input / $0.15 output / $0.004 cached input. Inception has not published an end date for the promotion, so we track the list price — the rate the model reverts to — rather than a discounted number that could expire without notice. Inception's own rate card shows both, listing Mercury 2.5 at $0.20 input / $0.02 cached / $0.75 output struck through, alongside the discounted prices. Inception's launch announcement states the same standard pricing of $0.20 / $0.75 per 1M tokens and describes $0.04 / $0.15 as a launch discount. Verified 2026-09-08. Verified against Google's Gemini API pricing page (ai.google.dev) on 2026-07-28: the Gemini 3.5 Flash-Lite Standard paid tier is $0.30/1M input, $2.50/1M output, $0.03/1M context caching (batch tier $0.15/$1.25). The identical $0.30/$2.50 price shared with the deprecated Gemini 2.5 Flash reflects Google's uniform Flash tier pricing, not a placeholder — the value matches LiteLLM gemini/gemini-3.5-flash-lite.
Specifications
Context window 260K tokens 1M tokens
Parameters Proprietary Proprietary
Speed (TPS)
Modalities
Input
text
textimageaudiovideo
Benchmarks sources: public leaderboards / Artificial Analysis, Google DeepMind Gemini 3.5 Flash-Lite model card, LMSYS Chatbot Arena (UC Berkeley)
Terminal-Bench 2.1 54 vendor
OSWorld-Verified 74 vendor
Independent composite scores each on its own scale — not part of the average above
AA Intelligence Index 22.7
AA Agentic Index 15.9
AA-Omniscience Index (−100–100) 5.2
GDPval-AA v2 1063.2 ELO
LMSYS Chatbot Arena 1456.9 ELO
Providers
Available from
Inception — $0.200/$0.750/Mtok
Google — $0.300/$2.50/Mtok

Cost at scale

1M tokens · 50/50 input/output
Projected cost of Mercury 2.5 vs Gemini 3.5 Flash-Lite at increasing token volumes
VolumeMercury 2.5Gemini 3.5 Flash-LiteSavings
1M tokens $0.34 $0.85 $0.51 (60%)
10M tokens $3.38 $8.5 $5.13 (60.4%)
100M tokens $33.75 $85 $51.25 (60.3%)
1000M tokens $337.5 $850 $512.5 (60.3%)

When to pick which

distilled from the pricing and spec data above
Pick Mercury 2.5 if…
  • cost dominates: $0.338/Mtok blended vs $0.850 — 60.3% less on the same 50/50 token mix
Pick Gemini 3.5 Flash-Lite if…
  • you need the longer context: 1M tokens vs 260K (3.8×)

Summary

Mercury 2.5 by Inception costs $0.200/Mtok input and $0.750/Mtok output, with a 260K-token context window. It supports text input.

Gemini 3.5 Flash-Lite by Google costs $0.300/Mtok input and $2.50/Mtok output, with a 1M-token context window. It supports text, image, audio, video input.

On a blended cost basis, Mercury 2.5 is 60.3% cheaper than Gemini 3.5 Flash-Lite.

Note: Pricing is per million tokens. Actual costs vary with usage patterns, prompt caching, and batch discounts. Always verify against official provider pricing pages.

More comparisons

pairs sharing a model with this page