Schematron V2 Small
by Inference.net · 3B parameters
Today's price · per 1M tokens
Input
$0.050
per 1M tokens
Output
$0.230
per 1M tokens
Blended
$0.095
blended $/1M — 3:1 weighted input:output
Cached input
$0.050
100% of input — prompt caching
How this price is scoped: Inference.net makes this model and is its only seller, so this is a first-party list price. Its own model catalogue at inference.net/models (read 2026-09-16, receipt data/evidence/inference-net-schematron-v2-small/2026-09-16.txt) prints the row "Schematron V2 Small / INFERENCE.NET / 125K input context / $0.05 / $0.05 / $0.23 / inference-net/schematron-v2-small" — the three money cells are input, cache read and output in that column order — and the model's own page at inference.net/models/schematron-v2-small prints the same three figures under "Model API pricing — Per 1M tokens". The catalogue's embedded record states them per token — costInputPerToken 0.000000050000, costCachedInputPerToken 0.000000050000, costOutputPerToken 0.000000230000 — which is $0.05 / $0.05 / $0.23 per 1M. OpenRouter's endpoint list for inference-net/schematron-v2-small (read 2026-09-16) returns exactly ONE endpoint, provider_name "InferenceNet", at prompt 0.00000005 / completion 0.00000023 / input_cache_read 0.00000005, so the maker's own figure and the only routed figure agree to the cent. No promotional or introductory marker appears against either SKU on the catalogue.
Price receipt
We read Inference.net's own pricing page on and found Schematron V2 Small listed with a price on the same row — that read is where the number above comes from, recorded as a new model. We keep the page text we read; its content hash is dcd2abfea5.
Overview
The accuracy tier of Inference.net's Schematron V2 pair — the same 3B HTML-to-JSON extraction job as Turbo, tuned for complex schemas and long pages rather than throughput, and priced above its sibling at $0.05/$0.23 per 1M tokens. Schema-constrained decoding, so the JSON conforms by construction. 128K input context, 4,096 max output tokens (half Turbo's), text only, no tool calling. Cache reads cost the same $0.05 as fresh input.
Capabilities
struck through = not supportedSpecifications
- Provider
- Inference.net
- Context window
- 128K tokens
- Modality
- text
- Parameters
- 3B
- Open source
- No — proprietary
- Released
- Status
- Current
- Last updated
- Tags
Availability verified: — listed on Inference.net's own page