ModelPriceWatch.com
Last scan 2026-09-15 Models tracked 259 Providers 32 Cheapest paid Granite 4.0 H Micro $0.017/Mtok in Every price links to its source

Inkling Small

by Thinking Machines · 276B total / 12B active parameters

Current budget open weights cheap tier

Today's price · per 1M tokens

Input

$0.300

per 1M tokens

Output

$1.20

per 1M tokens

Blended

$0.525

blended $/1M — 3:1 weighted input:output

Cached input

$0.060

20% of input — prompt caching

How this price is scoped: Tinker Serverless Inference (Beta) list price for the 256K-context nvfp4 sampling endpoint. OpenRouter lists this model's input at $0.50/1M — 67% over the vendor's own published rate — so the first-party figure is the one published here.

Source: official Thinking Machines pricing · read MODELPRICEWATCH.COM · 2026-09-15

Price receipt

We read Thinking Machines's own pricing page on and found Inkling-Small listed with a price on the same row — that read is where the number above comes from. We keep the page text we read; its content hash is 1e34d6686b.

Overview

The lighter Inkling: a 276B-total / 12B-active MoE that matches its larger sibling on many benchmarks at a quarter of the size, with the same native text/image/audio reasoning and variable thinking effort. Open weights on Hugging Face, sold through Tinker serverless inference at $0.30/$1.20 per 1M tokens.

Run it yourself

Deploy this open model on rented GPUs

Open weights mean you can self-host instead of paying per-token API prices. These platforms let you serve it on demand — often cheaper at scale.

Partner links — we may earn a commission if you sign up, at no cost to you. We only list platforms we'd recommend regardless.

Capabilities

struck through = not supported
Input 3/5
Text Image Audio Video PDF
Output 1/5
Text Image Audio Video Embedding
Features 5/9
Prompt caching Reasoning Coding Fast inference Long context Open weights Multimodal Web search Realtime

Benchmark performance

accuracy % · higher is better
Percentile vs all tracked models 40.6th3 independent measurements
Percentile per $/Mtok 78.1
GPQA Diamond vendor
88.3%
SWE-Bench Verified vendor
77.4%
Terminal-Bench 2.1 vendor
52.7%
MCP Atlas vendor
74.9%
Humanity's Last Exam vendor
29.6%

Every per-benchmark score we hold for this model is reported by its own vendor, so none is counted toward an accuracy average or any ranking. The percentile above is its standing across the independent composites below.

Independent composite scores
  • AA Intelligence Index: 26.1 (Artificial Analysis composite across reasoning, knowledge and coding evals)
  • AA-Omniscience Index: -8.9 (−100–100)
  • LMSYS Chatbot Arena: 1404.6 ELO (Human preference)
How it stacks up
  • Ranks #100 of 181 comparably-measured models by percentile score, across 3 independent measurements
  • Ranks #38 of 181 comparably-tested models by normalized performance per dollar

77.5 tokens/sec output

Source: Artificial Analysis, LMSYS Chatbot Arena (UC Berkeley), Thinking Machines launch post (Introducing Inkling) · updated · See full rankings →

Specifications

Context window
262K tokens
Modality
text, image, audio
Parameters
276B total / 12B active
Open source
Yes — open weights available
Released
Status
Current
Last updated
Tags
open-weightsbudgetreasoningmultimodalmoecaching

Availability verified: listed on Thinking Machines's own page