Cheapest LLM APIs by Token Cost
The cheapest LLM APIs ranked by token cost. Find budget-friendly models under $1/Mtok for high-volume production workloads.
55 models qualify
top 20 shown
sorted by input price, cheapest first
MODELPRICEWATCH.COM · 2026-09-22
Cost calculator for this use casemonthly cost, top 3 models
- 1 Granite 4.0 H Micro
- $—
- 2 Qwen3.7-Flash
- $—
- 3 Schematron V2 Turbo
- $—
Full ranking — top 20 models
list prices, USD per 1M tokens| Model | Blended* | Input | Output | Context | Provider |
|---|---|---|---|---|---|
| 1Granite 4.0 H Micro | $0.041 | $0.017 | $0.112 | 128K | IBM |
| 2Qwen3.7-Flash | $0.055 | $0.030 | $0.130 | 1M | Alibaba |
| 3Schematron V2 Turbo | $0.060 | $0.030 | $0.150 | 128K | Inference.net |
| 4LFM2.5 8B A1B | $0.052 | $0.030 | $0.120 | 33K | Together |
| 5GLM-OCR | $0.030 | $0.030 | $0.030 | 128K | Z.AI |
| 6Nova Micro | $0.061 | $0.035 | $0.140 | 128K | Amazon |
| 7GLM-4.6V-FlashX | $0.130 | $0.040 | $0.400 | 128K | Z.AI |
| 8Jev 1.13 | $0.032 | $0.042 | $0 | 64K | TypeSafe AI |
| 9Hy-MT2 1.8B | $0.077 | $0.044 | $0.177 | 8K | Tencent |
| 10Qwen-Flash | $0.138 | $0.050 | $0.400 | 1M | Alibaba |
| 11Qwen-Turbo | $0.088 | $0.050 | $0.200 | 1M | Alibaba |
| 12Schematron V2 Small | $0.095 | $0.050 | $0.230 | 128K | Inference.net |
| 13GPT-5 nano | $0.138 | $0.050 | $0.400 | 400K | OpenAI |
| 14Nova Lite | $0.105 | $0.060 | $0.240 | 300K | Amazon |
| 15Granite 4 H Small | $0.114 | $0.064 | $0.265 | 128K | IBM |
| 16Baichuan M2-32B | $0.070 | $0.070 | $0.070 | 131K | Baichuan |
| 17GPT-OSS 20B | $0.128 | $0.070 | $0.300 | 128K | Fireworks |
| 18GLM-4.7-FlashX | $0.153 | $0.070 | $0.400 | 128K | Z.AI |
| 19Hy-MT2 30B-A3B | $0.129 | $0.074 | $0.295 | 8K | Tencent |
| 20Qwen3-32B | $0.130 | $0.080 | $0.280 | 128K | DeepInfra |
Ranking basis: input price, cheapest first. * Blended = (3×input + 1×output) ÷ 4. Every price links to its source on the model page.
How models are selected
Budget-category models, sorted by input price per million tokens.
Prices are per million tokens (Mtok), sourced directly from — and linked to — official provider pricing pages. "Blended cost" is (3×input + 1×output) ÷ 4 — weighted toward input because real workloads read far more tokens than they generate.