Inference.net
Hosting provider · 2 models tracked
San Francisco company (Inference R&D, Inc.) running a serverless, OpenAI-compatible inference marketplace: it buys data centres' otherwise-idle GPU minutes and resells them, and its catalogue is mostly other makers' models. The rows tracked here are the exception — the Schematron V2 pair are Inference.net's own 3B HTML-to-JSON extraction models, sold first-party and by nobody else, at $0.03/1M input and $0.15/1M output (Turbo) and $0.05/1M input and $0.23/1M output (Small). Its founding year is not stated on any of its own pages and secondary sources disagree, so none is published here.
Inference.net pricing at a glanceSeptember 2026 · $ per 1M tokens
Inference.net API pricing (September 2026): 2 current models range from $0.030 to $0.050 per 1M input tokens and $0.150 to $0.230 per 1M output tokens. The cheapest paid model is Schematron V2 Turbo at $0.030/1M input; the priciest is Schematron V2 Small at $0.050/1M input / $0.230 output. Every price links to Inference.net's official pricing page and refreshes twice daily.
Today's Inference.net prices
2 models · cheapest blended firstSorted by blended cost (cheapest first). Prices per 1M tokens, September 2026 — every price links to Inference.net's official pricing page.
| Model | Blended* | Input | Output | Cached in | Relative cost | Context | Status |
|---|---|---|---|---|---|---|---|
| Schematron V2 Turbo | $0.060 |
$0.030 | $0.150 | $0.030 | 128K | Current | |
| Schematron V2 Small | $0.095 |
$0.050 | $0.230 | $0.050 | 128K | Current |
Quick stats
- Models tracked
- 2
- Type
- hosting
- Founded
- —
- Cheapest model
- Schematron V2 Turbo
- Cheapest blended
- $0.060/M