ModelPriceWatch.com
Last scan 2026-09-16 Models tracked 261 Providers 33 Cheapest paid Granite 4.0 H Micro $0.017/Mtok in Every price links to its source

Inference.net

Hosting provider · 2 models tracked

San Francisco company (Inference R&D, Inc.) running a serverless, OpenAI-compatible inference marketplace: it buys data centres' otherwise-idle GPU minutes and resells them, and its catalogue is mostly other makers' models. The rows tracked here are the exception — the Schematron V2 pair are Inference.net's own 3B HTML-to-JSON extraction models, sold first-party and by nobody else, at $0.03/1M input and $0.15/1M output (Turbo) and $0.05/1M input and $0.23/1M output (Small). Its founding year is not stated on any of its own pages and secondary sources disagree, so none is published here.

Inference.net pricing at a glanceSeptember 2026 · $ per 1M tokens

Inference.net API pricing (September 2026): 2 current models range from $0.030 to $0.050 per 1M input tokens and $0.150 to $0.230 per 1M output tokens. The cheapest paid model is Schematron V2 Turbo at $0.030/1M input; the priciest is Schematron V2 Small at $0.050/1M input / $0.230 output. Every price links to Inference.net's official pricing page and refreshes twice daily.

Models2
Input range$0.030–$0.050
Output range$0.150–$0.230
Cached-input tiers2

Today's Inference.net prices

2 models · cheapest blended first

Sorted by blended cost (cheapest first). Prices per 1M tokens, September 2026 — every price links to Inference.net's official pricing page.

Current Inference.net model prices per 1M tokens, sorted by blended cost
Model Blended* Input Output Cached in Relative cost Context Status
Schematron V2 Turbo
$0.060
$0.030 $0.150 $0.030
128K Current
Schematron V2 Small
$0.095
$0.050 $0.230 $0.050
128K Current
* Blended = (3×input + 1×output) ÷ 4 $/1M tokens · cheap · mid · expensive MODELPRICEWATCH.COM · 2026-09-16

Quick stats

Models tracked
2
Type
hosting
Founded
Cheapest model
Schematron V2 Turbo
Cheapest blended
$0.060/M