ModelPriceWatch.com
Last scan 2026-08-28 Models tracked 237 Providers 31 Cheapest paid Granite 4.0 H Micro $0.017/Mtok in Every price links to its source

DeepInfra

Hosting provider · 6 models tracked · Founded 2017

Serverless inference platform serving 100+ open-weight models (DeepSeek, Llama, Qwen, Gemma, Mistral) at pay-per-token rates — one of the cheapest per-token hosts for open models. No subscription or minimums.

DeepInfra pricing at a glanceAugust 2026 · $ per 1M tokens

DeepInfra API pricing (August 2026): 6 current models range from $0.080 to $1.30 per 1M input tokens and $0.180 to $2.60 per 1M output tokens. The cheapest paid model is Qwen3-32B at $0.080/1M input; the priciest is DeepSeek V4 Pro at $1.30/1M input / $2.60 output. Every price links to DeepInfra's official pricing page and refreshes twice daily.

Models6
Input range$0.080–$1.30
Output range$0.180–$2.60
Cached-input tiers3

Today's DeepInfra prices

6 models · cheapest blended first

Sorted by blended cost (cheapest first). Prices per 1M tokens, August 2026 — every price links to DeepInfra's official pricing page.

Current DeepInfra model prices per 1M tokens, sorted by blended cost
Model Blended* Input Output Cached in Relative cost Context Status
NVIDIA Nemotron 3.5 Lightning
$0.110
$0.080 $0.200 $0.040
256K Current
DeepSeek V4 Flash
$0.113
$0.090 $0.180 $0.018
1M Current
Qwen3-32B
$0.130
$0.080 $0.280
128K Current
Llama 4 Scout
$0.150
$0.100 $0.300
10M Current
Llama 4 Maverick
$0.350
$0.200 $0.800
1M Current
DeepSeek V4 Pro
$1.63
$1.30 $2.60 $0.100
1M Current
* Blended = (3×input + 1×output) ÷ 4 $/1M tokens · cheap · mid · expensive MODELPRICEWATCH.COM · 2026-08-28

Quick stats

Models tracked
6
Type
hosting
Founded
2017
Cheapest blended
$0.110/M

Open-weight models

4 open-weight models available from this provider.

Qwen3-32B
$0.080/M in · 128K ctx
Llama 4 Scout
$0.100/M in · 10M ctx
Llama 4 Maverick
$0.200/M in · 1M ctx
NVIDIA Nemotron 3.5 Lightning
$0.080/M in · 256K ctx
Run it yourself

Self-host DeepInfra's open-weight models

These models ship with open weights, so you can serve them yourself on rented GPUs instead of paying per-token API prices — often cheaper at scale.

Partner links — we may earn a commission if you sign up, at no cost to you. We only list platforms we'd recommend regardless.