Cheapest LLM APIs by Token Cost
The cheapest LLM APIs ranked by token cost. Find budget-friendly models under $1/Mtok for high-volume production workloads.
37 models qualify
top 20 shown
sorted by input price, cheapest first
MODELPRICEWATCH.COM · 2026-08-09
Cost calculator for this use casemonthly cost, top 3 models
- 1 Granite 4.0 Micro
- $—
- 2 Qwen3.7-Flash
- $—
- 3 GLM-OCR
- $—
Full ranking — top 20 models
list prices, USD per 1M tokens| Model | Blended* | Input | Output | Context | Provider |
|---|---|---|---|---|---|
| 1Granite 4.0 Micro | $0.041 | $0.017 | $0.112 | 128K | IBM |
| 2Qwen3.7-Flash | $0.055 | $0.030 | $0.130 | 1M | Alibaba |
| 3GLM-OCR | $0.030 | $0.030 | $0.030 | 128K | Z.AI |
| 4Nova Micro | $0.061 | $0.035 | $0.140 | 128K | Amazon |
| 5Qwen-Flash | $0.138 | $0.050 | $0.400 | 1M | Alibaba |
| 6Qwen-Turbo | $0.088 | $0.050 | $0.200 | 1M | Alibaba |
| 7Nova Lite | $0.105 | $0.060 | $0.240 | 300K | Amazon |
| 8Granite 4 H Small | $0.107 | $0.060 | $0.250 | 128K | IBM |
| 9Baichuan M2-32B | $0.070 | $0.070 | $0.070 | 33K | Baichuan |
| 10GPT OSS 20B | $0.128 | $0.070 | $0.300 | 128K | Fireworks |
| 11GLM-4.7-FlashX | $0.153 | $0.070 | $0.400 | 128K | Z.AI |
| 12Qwen3 32B | $0.130 | $0.080 | $0.280 | 128K | DeepInfra |
| 13DeepSeek V4 Flash | $0.113 | $0.090 | $0.180 | 1M | DeepInfra |
| 14Ministral 3 3B | $0.100 | $0.100 | $0.100 | 128K | Mistral |
| 15Voxtral Small 24B | $0.150 | $0.100 | $0.300 | 128K | Mistral |
| 16Reka Edge | $0.100 | $0.100 | $0.100 | 66K | Reka |
| 17GLM-4-32B-0414 | $0.100 | $0.100 | $0.100 | 128K | Z.AI |
| 18Hunyuan Hy3 Preview | $0.257 | $0.147 | $0.588 | 262K | Tencent |
| 19Command R 08-2024 | $0.262 | $0.150 | $0.600 | 128K | Cohere |
| 20GPT OSS 120B | $0.262 | $0.150 | $0.600 | 128K | Fireworks |
Ranking basis: input price, cheapest first. * Blended = (3×input + 1×output) ÷ 4. Every price links to its source on the model page.
How models are selected
Budget-category models, sorted by input price per million tokens.
Prices are per million tokens (Mtok), sourced directly from — and linked to — official provider pricing pages. "Blended cost" is (3×input + 1×output) ÷ 4 — weighted toward input because real workloads read far more tokens than they generate.