Schematron V2 Turbo
by Inference.net · 3B parameters
Today's price · per 1M tokens
Input
$0.030
per 1M tokens
Output
$0.150
per 1M tokens
Blended
$0.060
blended $/1M — 3:1 weighted input:output
Cached input
$0.030
100% of input — prompt caching
How this price is scoped: Inference.net makes this model and is its only seller, so this is a first-party list price. Its own model catalogue at inference.net/models (read 2026-09-16, receipt data/evidence/inference-net-schematron-v2-turbo/2026-09-16.txt) prints the row "Schematron V2 Turbo / INFERENCE.NET / 125K input context / $0.03 / $0.03 / $0.15 / inference-net/schematron-v2-turbo" — the three money cells are input, cache read and output in that column order — and the model's own page at inference.net/models/schematron-v2-turbo prints the same three figures under "Model API pricing — Per 1M tokens". The catalogue's embedded record states them per token — costInputPerToken 0.000000030000, costCachedInputPerToken 0.000000030000, costOutputPerToken 0.000000150000 — which is $0.03 / $0.03 / $0.15 per 1M. OpenRouter's endpoint list for inference-net/schematron-v2-turbo (read 2026-09-16) returns exactly ONE endpoint, provider_name "InferenceNet", at prompt 0.00000003 / completion 0.00000015 / input_cache_read 0.00000003 and discount 0, so the maker's own figure and the only routed figure agree to the cent. No promotional or introductory marker appears against either SKU on the catalogue.
Price receipt
We read Inference.net's own pricing page on and found Schematron V2 Turbo listed with a price on the same row — that read is where the number above comes from, recorded as a new model. We keep the page text we read; its content hash is dcd2abfea5.
Overview
Inference.net's throughput tier of Schematron V2, a 3B extraction model that turns raw HTML into schema-conforming JSON: you pass a JSON Schema in response_format instead of a prompt, and constrained decoding makes the output valid by construction rather than by retry. 128K input context, 8,192 max output tokens, text only, no tool calling. $0.03/$0.15 per 1M tokens. Cache reads are billed at the same $0.03 as fresh input, so caching buys latency here, not a discount.
Capabilities
struck through = not supportedSpecifications
- Provider
- Inference.net
- Context window
- 128K tokens
- Modality
- text
- Parameters
- 3B
- Open source
- No — proprietary
- Released
- Status
- Current
- Last updated
- Tags
Availability verified: — listed on Inference.net's own page