ModelPriceWatch.com
Last scan 2026-09-16 Models tracked 261 Providers 33 Cheapest paid Granite 4.0 H Micro $0.017/Mtok in Every price links to its source

Schematron V2 Small

by Inference.net · 3B parameters

Current new · 4d budget cheap tier
Availability: Served on Inference.net's own OpenAI-compatible API at api.inference.net/v1. Same two caveats as the Turbo row, read per SKU rather than assumed. The context window displays as "125K tokens" on both the catalogue and the model page while the record behind them says maxContextSize 128000; those are one number, because the site renders every window as floor(tokens/1024)K (1,000,000 shows as "977K", 200,000 as "195K"), so the published 128,000 is the maker's own figure. The catalogue's releaseDate field says 2026-04-16, which cannot be this SKU's launch — it repeats claude-opus-4-7's date on the same page and predates the SKU's first appearance in the OpenRouter catalogue on 2026-09-12 — so we publish the first provably purchasable date instead. Schematron V1 weights (Schematron-3B, Schematron-8B) are public on HuggingFace; the V2 repo inference-net/schematron-v2-llama-3.2-3b is not in the org's public model list and answers 401, so this row is open_source false.

Today's price · per 1M tokens

Input

$0.050

per 1M tokens

Output

$0.230

per 1M tokens

Blended

$0.095

blended $/1M — 3:1 weighted input:output

Cached input

$0.050

100% of input — prompt caching

How this price is scoped: Inference.net makes this model and is its only seller, so this is a first-party list price. Its own model catalogue at inference.net/models (read 2026-09-16, receipt data/evidence/inference-net-schematron-v2-small/2026-09-16.txt) prints the row "Schematron V2 Small / INFERENCE.NET / 125K input context / $0.05 / $0.05 / $0.23 / inference-net/schematron-v2-small" — the three money cells are input, cache read and output in that column order — and the model's own page at inference.net/models/schematron-v2-small prints the same three figures under "Model API pricing — Per 1M tokens". The catalogue's embedded record states them per token — costInputPerToken 0.000000050000, costCachedInputPerToken 0.000000050000, costOutputPerToken 0.000000230000 — which is $0.05 / $0.05 / $0.23 per 1M. OpenRouter's endpoint list for inference-net/schematron-v2-small (read 2026-09-16) returns exactly ONE endpoint, provider_name "InferenceNet", at prompt 0.00000005 / completion 0.00000023 / input_cache_read 0.00000005, so the maker's own figure and the only routed figure agree to the cent. No promotional or introductory marker appears against either SKU on the catalogue.

Source: official Inference.net pricing · read MODELPRICEWATCH.COM · 2026-09-16

Price receipt

We read Inference.net's own pricing page on and found Schematron V2 Small listed with a price on the same row — that read is where the number above comes from, recorded as a new model. We keep the page text we read; its content hash is dcd2abfea5.

Overview

The accuracy tier of Inference.net's Schematron V2 pair — the same 3B HTML-to-JSON extraction job as Turbo, tuned for complex schemas and long pages rather than throughput, and priced above its sibling at $0.05/$0.23 per 1M tokens. Schema-constrained decoding, so the JSON conforms by construction. 128K input context, 4,096 max output tokens (half Turbo's), text only, no tool calling. Cache reads cost the same $0.05 as fresh input.

Capabilities

struck through = not supported
Input 1/5
Text Image Audio Video PDF
Output 1/5
Text Image Audio Video Embedding
Features 2/9
Prompt caching Reasoning Coding Fast inference Long context Open weights Multimodal Web search Realtime

Specifications

Context window
128K tokens
Modality
text
Parameters
3B
Open source
No — proprietary
Released
Status
Current
Last updated
Tags
budget

Availability verified: listed on Inference.net's own page