Mercury 2.5 vs Gemini 3.5 Flash-Lite
Side-by-side comparison of API pricing, specs, benchmarks, and capabilities
Mercury 2.5 is 60.3% cheaper on blended cost ($0.338 vs $0.850/Mtok)
| Specification |
Mercury 2.5Intro price
by Inception
|
by Google
|
|---|---|---|
| Overview | ||
| Status | Current budget | Current budget |
| Released | ||
| Pricing per million tokens | ||
| Input | $0.200/Mtok | $0.300/Mtok |
| Output | $0.750/Mtok | $2.50/Mtok |
| Blended avg | $0.338/Mtok | $0.850/Mtok |
| Cached input | $0.020/Mtok | $0.030/Mtok |
| Price basis | These are Inception's LIST prices, per 1M tokens. Mercury 2.5 is currently on an 80%-off promotion, so the rate you actually pay today is $0.04 input / $0.15 output / $0.004 cached input. Inception has not published an end date for the promotion, so we track the list price — the rate the model reverts to — rather than a discounted number that could expire without notice. Inception's own rate card shows both, listing Mercury 2.5 at $0.20 input / $0.02 cached / $0.75 output struck through, alongside the discounted prices. Inception's launch announcement states the same standard pricing of $0.20 / $0.75 per 1M tokens and describes $0.04 / $0.15 as a launch discount. Verified 2026-09-08. | Verified against Google's Gemini API pricing page (ai.google.dev) on 2026-07-28: the Gemini 3.5 Flash-Lite Standard paid tier is $0.30/1M input, $2.50/1M output, $0.03/1M context caching (batch tier $0.15/$1.25). The identical $0.30/$2.50 price shared with the deprecated Gemini 2.5 Flash reflects Google's uniform Flash tier pricing, not a placeholder — the value matches LiteLLM gemini/gemini-3.5-flash-lite. |
| Specifications | ||
| Context window | 260K tokens | 1M tokens |
| Parameters | Proprietary | Proprietary |
| Speed (TPS) | — | — |
| Modalities | ||
| Input |
text
|
textimageaudiovideo
|
| Benchmarks sources: public leaderboards / Artificial Analysis, Google DeepMind Gemini 3.5 Flash-Lite model card, LMSYS Chatbot Arena (UC Berkeley) | ||
| Terminal-Bench 2.1 | — | 54 vendor |
| OSWorld-Verified | — | 74 vendor |
| Independent composite scores each on its own scale — not part of the average above | ||
| AA Intelligence Index | — | 22.7 |
| AA Agentic Index | — | 15.9 |
| AA-Omniscience Index (−100–100) | — | 5.2 |
| GDPval-AA v2 | — | 1063.2 ELO |
| LMSYS Chatbot Arena | — | 1456.9 ELO |
| Providers | ||
| Available from |
Inception — $0.200/$0.750/Mtok
|
Google — $0.300/$2.50/Mtok
|
Cost at scale
1M tokens · 50/50 input/output| Volume | Mercury 2.5 | Gemini 3.5 Flash-Lite | Savings |
|---|---|---|---|
| 1M tokens | $0.34 | $0.85 | $0.51 (60%) |
| 10M tokens | $3.38 | $8.5 | $5.13 (60.4%) |
| 100M tokens | $33.75 | $85 | $51.25 (60.3%) |
| 1000M tokens | $337.5 | $850 | $512.5 (60.3%) |
When to pick which
distilled from the pricing and spec data above
Pick Mercury 2.5 if…
- cost dominates: $0.338/Mtok blended vs $0.850 — 60.3% less on the same 50/50 token mix
Pick Gemini 3.5 Flash-Lite if…
- you need the longer context: 1M tokens vs 260K (3.8×)
Summary
Mercury 2.5 by Inception costs $0.200/Mtok input and $0.750/Mtok output, with a 260K-token context window. It supports text input.
Gemini 3.5 Flash-Lite by Google costs $0.300/Mtok input and $2.50/Mtok output, with a 1M-token context window. It supports text, image, audio, video input.
On a blended cost basis, Mercury 2.5 is 60.3% cheaper than Gemini 3.5 Flash-Lite.
Note: Pricing is per million tokens. Actual costs vary with usage patterns, prompt caching, and batch discounts. Always verify against official provider pricing pages.
More comparisons
pairs sharing a model with this page
Solar Pro 4 vs Gemini 3.5 Flash-Lite
94% price gap
Gemini 3.5 Flash-Lite vs Claude Haiku 4.5
58% price gap
Inkling Small vs Gemini 3.5 Flash-Lite
38% price gap
Mercury 2.5 vs GPT-5.6 Luna
25% price gap
DeepSeek V4 Flash Vision Exp vs Gemini 3.5 Flash-Lite
22% price gap
Gemini 3.5 Flash-Lite vs Gemini 2.5 Flash
GLM-4.7-Flash vs Ministral 3 3B
100% price gap
Solar Pro 4 vs DeepSeek V4 Flash
92% price gap