DeepInfra
Hosting provider · 6 models tracked · Founded 2017
Serverless inference platform serving 100+ open-weight models (DeepSeek, Llama, Qwen, Gemma, Mistral) at pay-per-token rates — one of the cheapest per-token hosts for open models. No subscription or minimums.
DeepInfra pricing at a glanceAugust 2026 · $ per 1M tokens
DeepInfra API pricing (August 2026): 6 current models range from $0.080 to $1.30 per 1M input tokens and $0.180 to $2.60 per 1M output tokens. The cheapest paid model is Qwen3-32B at $0.080/1M input; the priciest is DeepSeek V4 Pro at $1.30/1M input / $2.60 output. Every price links to DeepInfra's official pricing page and refreshes twice daily.
Today's DeepInfra prices
6 models · cheapest blended firstSorted by blended cost (cheapest first). Prices per 1M tokens, August 2026 — every price links to DeepInfra's official pricing page.
| Model | Blended* | Input | Output | Cached in | Relative cost | Context | Status |
|---|---|---|---|---|---|---|---|
| NVIDIA Nemotron 3.5 Lightning | $0.110 |
$0.080 | $0.200 | $0.040 | 256K | Current | |
| DeepSeek V4 Flash | $0.113 |
$0.090 | $0.180 | $0.018 | 1M | Current | |
| Qwen3-32B | $0.130 |
$0.080 | $0.280 | — | 128K | Current | |
| Llama 4 Scout | $0.150 |
$0.100 | $0.300 | — | 10M | Current | |
| Llama 4 Maverick | $0.350 |
$0.200 | $0.800 | — | 1M | Current | |
| DeepSeek V4 Pro | $1.63 |
$1.30 | $2.60 | $0.100 | 1M | Current |
Quick stats
- Models tracked
- 6
- Type
- hosting
- Founded
- 2017
- Cheapest model
- NVIDIA Nemotron 3.5 Lightning
- Cheapest blended
- $0.110/M
Open-weight models
4 open-weight models available from this provider.
Self-host DeepInfra's open-weight models
These models ship with open weights, so you can serve them yourself on rented GPUs instead of paying per-token API prices — often cheaper at scale.
Partner links — we may earn a commission if you sign up, at no cost to you. We only list platforms we'd recommend regardless.