ModelPriceWatch.com
Last scan 2026-09-15 Models tracked 259 Providers 32 Cheapest paid Granite 4.0 H Micro $0.017/Mtok in Every price links to its source

Gemini 3.5 Flash

by Google

Current flagship mid tier

Today's price · per 1M tokens

Input

$1.50

per 1M tokens

Output

$9.00

per 1M tokens

Blended

$3.38

blended $/1M — 3:1 weighted input:output

Source: official Google pricing · read MODELPRICEWATCH.COM · 2026-09-15

Price receipt

We read Google's own pricing page on and found Gemini 3.5 Flash listed with a price on the same card — that read is where the number above comes from. We keep the page text we read; its content hash is 2b5d2335c5.

Overview

Google's fast, low-cost workhorse from I/O 2026 — strong multimodal reasoning and coding (beats Gemini 3.1 Pro), with native computer-use for agents. The default Gemini across Search and the API. $1.50/$9 per 1M tokens.

Capabilities

struck through = not supported
Input 4/5
Text Image Audio Video PDF
Output 1/5
Text Image Audio Video Embedding
Features 3/9
Prompt caching Reasoning Coding Fast inference Long context Open weights Multimodal Web search Realtime

Benchmark performance

accuracy % · higher is better
Avg benchmark score70.1%
Perf per $/Mtok20.8
GPQA Diamond vendor
72%
SWE-Bench Verified vendor
55%
Terminal-Bench 2.1
76.2%
MCP Atlas
83.6%
OSWorld-Verified
78.4%
Humanity's Last Exam
40.2%
ARC-AGI 2
72.1%
AIME 2025 vendor
78%
MMMLU vendor
85%
BFCL vendor
75%
HumanEval vendor
88%
MATH 500 vendor
82%
Independent composite scores
  • AA Intelligence Index: 33 (Artificial Analysis composite across reasoning, knowledge and coding evals)
  • AA-Omniscience Index: 21.2 (−100–100)
  • LMSYS Chatbot Arena: 1477.7 ELO (Human preference)
How it stacks up
  • Ranks #37 of 84 benchmarked models by average score
  • Ranks #31 of 181 comparably-measured models by percentile score, across 8 independent measurements
  • Ranks #124 of 181 comparably-tested models by normalized performance per dollar
  • Strongest at MCP Atlas — 83.6%, #2 of 28

175.4 tokens/sec output 23.16s latency to first token (TTFT)

Source: Vellum LLM Leaderboard · updated · See full rankings →

Specifications

Provider
Google
Context window
1M tokens
Modality
text, image, audio, video
Parameters
Proprietary
Open source
No — proprietary
Released
Status
Current
Last updated
Tags
fastmultimodalflagship

Availability verified: listed on Google's own page