Deprecated — shuts down Oct 16, 2026.
Still purchasable until then, but the provider has
announced end-of-life. Don't start new work on it.
Availability: Earliest possible retirement date; successor gemini-3.1-flash-lite.
Today's price · per 1M tokens
Input
$0.100
per 1M tokens
Output
$0.400
per 1M tokens
Blended
$0.250
avg of input & output $/1M
Source: official Google pricing · read Jun 25, 2026
MODELPRICEWATCH.COM · 2026-08-01
Overview
Lightweight, low-cost model. $0.10/$0.40 per 1M tokens.
Capabilities
struck through = not supportedInput 4/5
Text ✓
Image ✓
Audio ✓
Video ✓
PDF
Output 1/5
Text ✓
Image
Audio
Video
Embedding
Features 3/9
Prompt caching
Reasoning
Coding
Fast inference ✓
Long context ✓
Open weights
Multimodal ✓
Web search
Realtime
Benchmark performance
accuracy % · higher is betterGPQA Diamond vendor
40%
SWE-Bench Verified vendor
22%
Humanity's Last Exam vendor
10%
ARC-AGI 2 vendor
10%
AIME 2025 vendor
42%
MMMLU vendor
70%
BFCL vendor
55%
HumanEval vendor
75%
MATH 500 vendor
58%
Every per-benchmark score we hold for this model is reported by its own vendor, so none is counted toward an accuracy average or any ranking.
450 tokens/sec output
Source: Google model card · updated Feb 1, 2026 · See full rankings →
Specifications
- Provider
- Context window
- 1M tokens
- Modality
- text, image, audio, video
- Parameters
- Proprietary
- Open source
- No — proprietary
- Released
- Jun 17, 2025
- Status
- Deprecated ends Oct 16, 2026
- Last updated
- Jun 25, 2026
- Tags
Availability verified: Jul 25, 2026 — per Google's own deprecation notice