ModelPriceWatch.com
Last scan 2026-08-01 Models tracked 187 Providers 28 Cheapest paid Granite 4.0 Micro $0.017/Mtok in Every price links to its source

Inkling Small vs Gemini 3.5 Flash-Lite

Side-by-side comparison of API pricing, specs, benchmarks, and capabilities

Inkling Small is 38.2% cheaper on blended cost ($0.525 vs $0.850/Mtok)
Specification
Inkling SmallIntro price
Overview
StatusCurrent budget Open weights Current budget
Released Jul 30, 2026 Jul 21, 2026
Pricing per million tokens
Input $0.300/Mtok $0.300/Mtok
Output $1.20/Mtok $2.50/Mtok
Blended avg $0.525/Mtok $0.850/Mtok
Cached input $0.060/Mtok $0.030/Mtok
Specifications
Context window 262K tokens 1M tokens
Parameters 276B total / 12B active Proprietary
Speed (TPS)
Modalities
Input
textimageaudio
textimageaudiovideo
Benchmarks source: Artificial Analysis
Independent composite scores measured for both models — each on its own scale
AA Intelligence Index 40.2 36.5
AA Agentic Index 30.8 26.8
AA-Omniscience Index (−100–100) -9 6.9
LMSYS Chatbot Arena 1430.5 ELO 1456.5 ELO
GDPval-AA v2 1138 ELO
Providers
Available from
Thinking Machines — $0.300/$1.20/Mtok
Google — $0.300/$2.50/Mtok

Cost at scale

1M tokens · 50/50 input/output
Projected cost of Inkling Small vs Gemini 3.5 Flash-Lite at increasing token volumes
VolumeInkling SmallGemini 3.5 Flash-LiteSavings
1M tokens $0.52 $0.85 $0.33 (38.8%)
10M tokens $5.25 $8.5 $3.25 (38.2%)
100M tokens $52.5 $85 $32.5 (38.2%)
1000M tokens $525 $850 $325 (38.2%)

When to pick which

distilled from the pricing and spec data above
Pick Inkling Small if…
  • cost dominates: $0.525/Mtok blended vs $0.850 — 38.2% less on the same 50/50 token mix
  • you want open weights — self-host it, fine-tune it, or exit the API entirely; Gemini 3.5 Flash-Lite is closed
Pick Gemini 3.5 Flash-Lite if…
  • your workload re-reads context (agents, RAG, long chats): cached input costs $0.030/Mtok — 90% off its list input price
  • you need the longer context: 1M tokens vs 262K (3.8×)

Summary

Inkling Small by Thinking Machines costs $0.300/Mtok input and $1.20/Mtok output, with a 262K-token context window. It supports text, image, audio input.

Gemini 3.5 Flash-Lite by Google costs $0.300/Mtok input and $2.50/Mtok output, with a 1M-token context window. It supports text, image, audio, video input.

On a blended cost basis, Inkling Small is 38.2% cheaper than Gemini 3.5 Flash-Lite.

Note: Pricing is per million tokens. Actual costs vary with usage patterns, prompt caching, and batch discounts. Always verify against official provider pricing pages.

More comparisons

pairs sharing a model with this page