ModelPriceWatch.com
Last scan 2026-09-16 Models tracked 261 Providers 33 Cheapest paid Granite 4.0 H Micro $0.017/Mtok in Every price links to its source

Schematron V2 Turbo

by Inference.net · 3B parameters

Current new · 4d budget cheap tier
Availability: Served on Inference.net's own OpenAI-compatible API at api.inference.net/v1; the served build is tagged inference-net/schematron-v2-turbo-20260902. Two caveats a reader may hit. First, the context window: the catalogue and the model page both display "125K tokens", while the record behind them says maxContextSize 128000. These are the same number — the site renders every window as floor(tokens/1024)K, which is why 1,000,000 shows as "977K" and 200,000 as "195K" on the same page — so the published 128,000 is the maker's figure, not a third-party one. Second, the release date: the catalogue record carries releaseDate 2026-04-16, which cannot be this SKU's launch (the served build is dated 2026-09-02, the SKU first appears in the OpenRouter catalogue on 2026-09-12, and the field repeats claude-opus-4-7's date on the same page), so we publish the first date on which it is provably purchasable rather than the maker's field. Schematron V1 (Schematron-3B, Schematron-8B) has public weights on HuggingFace; V2 does not — inference-net/schematron-v2-granite-4.0-h-micro is not in the org's public model list and answers 401 — so this row is open_source false.

Today's price · per 1M tokens

Input

$0.030

per 1M tokens

Output

$0.150

per 1M tokens

Blended

$0.060

blended $/1M — 3:1 weighted input:output

Cached input

$0.030

100% of input — prompt caching

How this price is scoped: Inference.net makes this model and is its only seller, so this is a first-party list price. Its own model catalogue at inference.net/models (read 2026-09-16, receipt data/evidence/inference-net-schematron-v2-turbo/2026-09-16.txt) prints the row "Schematron V2 Turbo / INFERENCE.NET / 125K input context / $0.03 / $0.03 / $0.15 / inference-net/schematron-v2-turbo" — the three money cells are input, cache read and output in that column order — and the model's own page at inference.net/models/schematron-v2-turbo prints the same three figures under "Model API pricing — Per 1M tokens". The catalogue's embedded record states them per token — costInputPerToken 0.000000030000, costCachedInputPerToken 0.000000030000, costOutputPerToken 0.000000150000 — which is $0.03 / $0.03 / $0.15 per 1M. OpenRouter's endpoint list for inference-net/schematron-v2-turbo (read 2026-09-16) returns exactly ONE endpoint, provider_name "InferenceNet", at prompt 0.00000003 / completion 0.00000015 / input_cache_read 0.00000003 and discount 0, so the maker's own figure and the only routed figure agree to the cent. No promotional or introductory marker appears against either SKU on the catalogue.

Source: official Inference.net pricing · read MODELPRICEWATCH.COM · 2026-09-16

Price receipt

We read Inference.net's own pricing page on and found Schematron V2 Turbo listed with a price on the same row — that read is where the number above comes from, recorded as a new model. We keep the page text we read; its content hash is dcd2abfea5.

Overview

Inference.net's throughput tier of Schematron V2, a 3B extraction model that turns raw HTML into schema-conforming JSON: you pass a JSON Schema in response_format instead of a prompt, and constrained decoding makes the output valid by construction rather than by retry. 128K input context, 8,192 max output tokens, text only, no tool calling. $0.03/$0.15 per 1M tokens. Cache reads are billed at the same $0.03 as fresh input, so caching buys latency here, not a discount.

Capabilities

struck through = not supported
Input 1/5
Text Image Audio Video PDF
Output 1/5
Text Image Audio Video Embedding
Features 3/9
Prompt caching Reasoning Coding Fast inference Long context Open weights Multimodal Web search Realtime

Specifications

Context window
128K tokens
Modality
text
Parameters
3B
Open source
No — proprietary
Released
Status
Current
Last updated
Tags
budgetfast

Availability verified: listed on Inference.net's own page