ModelPriceWatch.com
Last scan 2026-08-01 Models tracked 187 Providers 28 Cheapest paid Granite 4.0 Micro $0.017/Mtok in Every price links to its source

Inkling

by Thinking Machines · 975B total / 41B active parameters

Current new · 17d Intro price mid tier open weights mid tier

Today's price · per 1M tokens

Input

$1.00

per 1M tokens

Output

$4.05

per 1M tokens

Blended

$1.76

blended $/1M — 3:1 weighted input:output

Cached input

$0.170

17% of input — prompt caching

Source: official Thinking Machines pricing · read Aug 1, 2026 MODELPRICEWATCH.COM · 2026-08-01

Overview

Thinking Machines' first frontier model: a 975B-total / 41B-active Mixture-of-Experts transformer that reasons natively over text, images and audio, with controllable thinking effort. The weights are open on Hugging Face, and the lab also sells it directly through Tinker's serverless inference at $1.00/$4.05 per 1M tokens, with cached input at $0.17.

Run it yourself

Deploy this open model on rented GPUs

Open weights mean you can self-host instead of paying per-token API prices. These platforms let you serve it on demand — often cheaper at scale.

Partner links — we may earn a commission if you sign up, at no cost to you. We only list platforms we'd recommend regardless.

Capabilities

struck through = not supported
Input 3/5
Text ✓ Image ✓ Audio ✓ Video PDF
Output 1/5
Text ✓ Image Audio Video Embedding
Features 5/9
Prompt caching ✓ Reasoning ✓ Coding Fast inference Long context ✓ Open weights ✓ Multimodal ✓ Web search Realtime

Benchmark performance

accuracy % · higher is better
Percentile vs all tracked models 50.7th5 independent measurements
Percentile per $/Mtok 28.8

We hold no per-benchmark accuracy scores for this model yet, so it has no accuracy average. That is a gap in our coverage, not a sign the model is untested — independent evaluators often publish a composite index for a new model long before releasing its per-benchmark numbers. The percentile above is its standing across the independent composites below.

Independent composite scores
  • AA Intelligence Index: 40.7 (Artificial Analysis composite across reasoning, knowledge and coding evals)
  • AA Agentic Index: 32.3 (Tool use, planning, autonomy)
  • AA-Omniscience Index: 2 (−100–100)
  • GDPval-AA v2: 1236.8 ELO (Real-world work tasks, human baseline = 1000)
  • LMSYS Chatbot Arena: 1440.8 ELO (Human preference)
How it stacks up
  • Ranks #51 of 136 comparably-measured models by percentile score, across 5 independent measurements
  • Ranks #73 of 136 comparably-tested models by normalized performance per dollar

Source: Artificial Analysis · updated Aug 1, 2026 · See full rankings →

Specifications

Context window
262K tokens
Modality
text, image, audio
Parameters
975B total / 41B active
Open source
Yes — open weights available
Released
Jul 15, 2026
Status
Current
Last updated
Aug 1, 2026
Tags
open-weightsmid-tierreasoningmultimodalmoecaching

Availability verified: Aug 1, 2026 — listed on Thinking Machines's own page