Today's price · per 1M tokens
Input
$2.00
per 1M tokens
Output
$12.00
per 1M tokens
Blended
$7.00
avg of input & output $/1M
Cached input
$0.200
10% of input — prompt caching
Source: official Google pricing · read Jul 4, 2026
MODELPRICEWATCH.COM · 2026-08-01
Overview
Frontier reasoning model. $2/$12 per 1M (≤200K context), 2M context window.
Capabilities
struck through = not supportedInput 4/5
Text ✓
Image ✓
Audio ✓
Video ✓
PDF
Output 1/5
Text ✓
Image
Audio
Video
Embedding
Features 4/9
Prompt caching ✓
Reasoning ✓
Coding
Fast inference
Long context ✓
Open weights
Multimodal ✓
Web search
Realtime
Benchmark performance
accuracy % · higher is betterAvg benchmark score74.8%
Perf per $/Mtok10.7
GPQA Diamond
94.3%
Terminal-Bench 2.1
70.3%
MCP Atlas
69.2%
BrowseComp
85.9%
OSWorld-Verified
76.2%
Humanity's Last Exam
44.4%
ARC-AGI 2
77.1%
HumanEval
80.6%
Independent composite scores
- AA Intelligence Index: 46.5 (Artificial Analysis composite across reasoning, knowledge and coding evals)
- AA Agentic Index: 21.4 (Tool use, planning, autonomy)
- AA-Omniscience Index: 32.9 (−100–100)
- LMSYS Chatbot Arena: 1485.4 ELO (Human preference)
How it stacks up
- Ranks #26 of 68 benchmarked models by average score
- Ranks #26 of 133 comparably-measured models by percentile score, across 12 independent measurements
- Ranks #100 of 133 comparably-tested models by normalized performance per dollar
- Strongest at GPQA Diamond — 94.3%, #3 of 57
136.2 tokens/sec output 20.34s latency to first token (TTFT)
Source: Vellum LLM Leaderboard · updated Aug 1, 2026 · See full rankings →
Specifications
- Provider
- Context window
- 2M tokens
- Modality
- text, image, audio, video
- Parameters
- Proprietary
- Open source
- No — proprietary
- Released
- Nov 1, 2025
- Status
- Current
- Last updated
- Jul 4, 2026
- Tags
Availability verified: Jul 20, 2026 — listed on Google's own page