ModelPriceWatch.com
Last scan 2026-09-04 Models tracked 244 Providers 31 Cheapest paid Granite 4.0 H Micro $0.017/Mtok in Every price links to its source

Gemini 3.5 Flash-Lite

by Google

Current new · 45d budget cheap tier

Today's price · per 1M tokens

Input

$0.300

per 1M tokens

Output

$2.50

per 1M tokens

Blended

$0.850

blended $/1M — 3:1 weighted input:output

Cached input

$0.030

10% of input — prompt caching

How this price is scoped: Verified against Google's Gemini API pricing page (ai.google.dev) on 2026-07-28: the Gemini 3.5 Flash-Lite Standard paid tier is $0.30/1M input, $2.50/1M output, $0.03/1M context caching (batch tier $0.15/$1.25). The identical $0.30/$2.50 price shared with the deprecated Gemini 2.5 Flash reflects Google's uniform Flash tier pricing, not a placeholder — the value matches LiteLLM gemini/gemini-3.5-flash-lite.

Source: official Google pricing · read MODELPRICEWATCH.COM · 2026-09-04

Price receipt

We read Google's own pricing page on and found Gemini 3.5 Flash-Lite listed with a price on the same card — that read is where the number above comes from. We keep the page text we read; its content hash is 2b5d2335c5.

Overview

Ultra-cheap, fast tier of the Gemini 3.5 family at $0.30/1M input and $2.50/1M output, with full multimodal input and a 1M-token context. Cheapest way to run agentic and high-volume workloads on Gemini 3.5.

Capabilities

struck through = not supported
Input 4/5
Text Image Audio Video PDF
Output 1/5
Text Image Audio Video Embedding
Features 4/9
Prompt caching Reasoning Coding Fast inference Long context Open weights Multimodal Web search Realtime

Benchmark performance

accuracy % · higher is better
Percentile vs all tracked models 46th5 independent measurements
Percentile per $/Mtok 54.1
Terminal-Bench 2.1 vendor
54%
OSWorld-Verified vendor
74%

Every per-benchmark score we hold for this model is reported by its own vendor, so none is counted toward an accuracy average or any ranking. The percentile above is its standing across the independent composites below.

Independent composite scores
  • AA Intelligence Index: 37.4 (Artificial Analysis composite across reasoning, knowledge and coding evals)
  • AA Agentic Index: 27.2 (Tool use, planning, autonomy)
  • AA-Omniscience Index: 5.2 (−100–100)
  • GDPval-AA v2: 1136.3 ELO (Real-world work tasks, human baseline = 1000)
  • LMSYS Chatbot Arena: 1456.9 ELO (Human preference)
How it stacks up
  • Ranks #84 of 176 comparably-measured models by percentile score, across 5 independent measurements
  • Ranks #46 of 176 comparably-tested models by normalized performance per dollar

365.2 tokens/sec output

Source: Artificial Analysis, Google DeepMind Gemini 3.5 Flash-Lite model card, LMSYS Chatbot Arena (UC Berkeley) · updated · See full rankings →

Specifications

Provider
Google
Context window
1M tokens
Modality
text, image, audio, video
Parameters
Proprietary
Open source
No — proprietary
Released
Status
Current
Last updated
Tags
fastmultimodalbudget

Availability verified: listed on Google's own page