Deprecated — shuts down Oct 16, 2026.
Still purchasable until then, but the provider has
announced end-of-life. Don't start new work on it.
Availability: Earliest possible retirement date; successor gemini-3.6-flash.
Today's price · per 1M tokens
Input
$0.300
per 1M tokens
Output
$2.50
per 1M tokens
Blended
$1.40
avg of input & output $/1M
Source: official Google pricing · read Jul 4, 2026
MODELPRICEWATCH.COM · 2026-08-01
Overview
One of the cheapest hosted APIs. $0.075/$0.30 per 1M tokens, fast and multimodal.
Capabilities
struck through = not supportedInput 4/5
Text ✓
Image ✓
Audio ✓
Video ✓
PDF
Output 1/5
Text ✓
Image
Audio
Video
Embedding
Features 3/9
Prompt caching
Reasoning
Coding
Fast inference ✓
Long context ✓
Open weights
Multimodal ✓
Web search
Realtime
Benchmark performance
accuracy % · higher is betterAvg benchmark score50.2%
Perf per $/Mtok35.9
GPQA Diamond
78.3%
SWE-Bench Verified
38%
Terminal-Bench 2.1
16.9%
LiveCodeBench
63.5%
Humanity's Last Exam
12.1%
ARC-AGI 2
15%
AIME 2025
78%
MMMLU
78%
MATH 500
72%
Independent composite scores
- LMSYS Chatbot Arena: 1410.2 ELO (Human preference)
How it stacks up
- Ranks #57 of 68 benchmarked models by average score
- Ranks #82 of 133 comparably-measured models by percentile score, across 10 independent measurements
- Ranks #65 of 133 comparably-tested models by normalized performance per dollar
- Strongest at MATH 500 — 72%, #10 of 23
200 tokens/sec output 0.35s latency to first token (TTFT)
Source: Vellum LLM Leaderboard · updated Aug 1, 2026 · See full rankings →
Specifications
- Provider
- Context window
- 1M tokens
- Modality
- text, image, audio, video
- Parameters
- Proprietary
- Open source
- No — proprietary
- Released
- Jun 17, 2025
- Status
- Deprecated ends Oct 16, 2026
- Last updated
- Jul 4, 2026
- Tags
Availability verified: Jul 25, 2026 — per Google's own deprecation notice