Open-weight models
70 open-weight models hosted across inference providers. Compare hosting prices — the model weights are free, you pay only for compute.
Open-weight models
70
current, verified hosted prices
Cheapest input /Mtok
$0
GLM-4.7-Flash
Fastest inference
1000 TPS
GPT-OSS 20B
Largest context
10M tokens
Llama 4 Scout
Today's open-weight prices
70 models · sorted by input price, cheapest first| Parameters | Notes | ||||
|---|---|---|---|---|---|
| GLM-4.7-FlashZ.AI | $0 | $0 | 128K | — | |
| Granite 4.0 MicroIBM | $0.017 | $0.112 | 128K | — | |
| GLM-OCRZ.AI | $0.030 | $0.030 | 128K | — | |
| Llama 3.1 8B InstantGroq | $0.050 | $0.080 | 128K | 8B | 840 TPS |
| Llama 3.1 8BMeta | $0.050 | $0.080 | 128K | 8B | |
| Granite 4 H SmallIBM | $0.060 | $0.250 | 128K | — | |
| Baichuan M2-32BBaichuan | $0.070 | $0.070 | 33K | 32B | |
| GPT OSS 20BFireworks | $0.070cached $0.035 | $0.300 | 128K | 20B | |
| GLM-4.7-FlashXZ.AI | $0.070cached $0.010 | $0.400 | 128K | — | |
| GPT-OSS 20BGroq | $0.075 | $0.300 | 128K | 20B | 1000 TPS |
| Qwen3 32BDeepInfra | $0.080 | $0.280 | 128K | 32B | |
| Ministral 3 3BMistral | $0.100 | $0.100 | 128K | 3B | |
| Voxtral Small 24BMistral | $0.100 | $0.300 | 128K | 24B | |
| GLM-4-32B-0414Z.AI | $0.100 | $0.100 | 128K | 32B | |
| Llama 4 ScoutDeepInfra | $0.100 | $0.300 | 10M | 17B (16 experts) | |
| Granite Embedding 278M MultilingualIBM | $0.106 | $— | — | 278M | |
| Llama 4 ScoutMeta | $0.110 | $0.340 | 10M | 17B (16 experts) | |
| GPT OSS 120BFireworks | $0.150cached $0.015 | $0.600 | 128K | 120B | |
| GPT-OSS 120BGroq | $0.150 | $0.600 | 128K | 120B | 500 TPS |
| Granite 4 H MediumIBM | $0.150 | $0.600 | 128K | — | |
| gpt-oss-120BTogether | $0.150 | $0.600 | 128K | 120B | |
| Qwen3-32BAlibaba | $0.160 | $0.640 | 131K | 32B | |
| Jamba MiniAI21 Labs | $0.200 | $0.400 | 256K | — | |
| GLM-4.5-AirZ.AI | $0.200cached $0.030 | $1.10 | 128K | — | |
| Llama 4 MaverickDeepInfra | $0.200 | $0.800 | 1M | 17B (128 experts) | |
| Gemma-4-31B-it-PearlTogether | $0.280 | $0.860 | 128K | 31B | |
| Qwen3 32BGroq | $0.290 | $0.590 | 128K | 32B | 662 TPS |
| MiniMax 2.5Fireworks | $0.300cached $0.030 | $1.20 | 128K | — | |
| MiniMax 2.7Fireworks | $0.300cached $0.060 | $1.20 | 128K | — | |
| MiniMax M3Fireworks | $0.300cached $0.060 | $1.20 | 1M | — | |
| Granite 4 H LargeIBM | $0.300 | $1.20 | 128K | — | |
| MiniMax-M2.7MiniMax | $0.300cached $0.060 | $1.20 | 205K | — | |
| MiniMax-M3MiniMax | $0.300cached $0.060 | $1.20 | 1M | — | |
| MiniMax M3Together | $0.300cached $0.060 | $1.20 | 1M | — | |
| GLM-4.6VZ.AI | $0.300cached $0.050 | $0.900 | 128K | — | |
| Inkling SmallThinking Machines | $0.300cached $0.060 | $1.20 | 262K | 276B total / 12B active | |
| Qwen3.7-PlusTogether | $0.320 | $1.28 | 128K | — | |
| Qwen 3.7 PlusFireworks | $0.400cached $0.080 | $1.60 | 128K | — | |
| Qwen 3.6 PlusFireworks | $0.500cached $0.100 | $3.00 | 128K | — | |
| Llama 3.3 70B VersatileGroq | $0.590 | $0.790 | 128K | 70B | 394 TPS |
| Llama 3.3 70BMeta | $0.590 | $0.790 | 128K | 70B | |
| Qwen3.6-27BAlibaba | $0.600 | $3.60 | 262K | 27B | |
| Qwen3.5-397B-A17BAlibaba | $0.600 | $3.60 | 262K | 397B (A17B) | |
| Kimi K2.5Fireworks | $0.600cached $0.100 | $3.00 | 256K | — | |
| NVIDIA Nemotron 3 UltraFireworks | $0.600cached $0.120 | $2.40 | 128K | — | |
| Qwen 3.6 27BGroq | $0.600 | $3.00 | 128K | 27B | 500 TPS |
| Kimi K2.5Moonshot | $0.600cached $0.100 | $3.00 | 262K | — | |
| NVIDIA Nemotron 3 UltraTogether | $0.600cached $0.200 | $3.60 | 128K | — | |
| GLM-4.5Z.AI | $0.600cached $0.110 | $2.20 | 128K | — | |
| GLM-4.6Z.AI | $0.600cached $0.110 | $2.20 | 200K | — | |
| GLM-4.7Z.AI | $0.600cached $0.110 | $2.20 | 128K | — | |
| LongCat-2.0Meituan | $0.750cached $0.015 | $2.95 | 262K | — | |
| Kimi K2.6Fireworks | $0.950cached $0.160 | $4.00 | 256K | — | |
| Kimi K2.7 CodeFireworks | $0.950cached $0.190 | $4.00 | 256K | — | |
| Kimi K2.6Moonshot | $0.950cached $0.160 | $4.00 | 262K | — | |
| Kimi K2.7 CodeMoonshot | $0.950cached $0.190 | $4.00 | 262K | — | |
| Kimi K2.7 CodeTogether | $0.950cached $0.190 | $4.00 | 256K | — | |
| GLM-5Z.AI | $1.00cached $0.200 | $3.20 | 128K | — | |
| InklingThinking Machines | $1.00cached $0.170 | $4.05 | 262K | 975B total / 41B active | |
| Llama 3.3 70BTogether | $1.04 | $1.04 | 128K | 70B | |
| GLM-5-TurboZ.AI | $1.20cached $0.240 | $4.00 | 128K | — | |
| GLM-5V-TurboZ.AI | $1.20cached $0.240 | $4.00 | 128K | — | |
| GLM 5.1Fireworks | $1.40cached $0.260 | $4.40 | 128K | — | |
| GLM 5.2Fireworks | $1.40cached $0.140 | $4.40 | 1M | — | |
| GLM-5.2Together | $1.40cached $0.260 | $4.40 | 1M | — | |
| GLM-5.1Z.AI | $1.40cached $0.260 | $4.40 | 128K | — | |
| GLM-5.2Z.AI | $1.40cached $0.260 | $4.40 | 1M | — | |
| Kimi K2.7 Code HighSpeedMoonshot | $1.90cached $0.380 | $8.00 | 262K | — | |
| Jamba LargeAI21 Labs | $2.00 | $8.00 | 256K | — | |
| Kimi K3Moonshot | $3.00cached $0.300 | $15.00 | 1M | — |
MODELPRICEWATCH.COM · 2026-08-09
How open-weight pricing works
Open weight models have freely available model weights — anyone can download and run them. You pay only for the compute to serve inference. Hosting providers like Groq (custom LPU chips, fastest inference), Together AI (200+ models, competitive pricing), and Fireworks offer per-token pricing without lock-in.
The same model can have very different prices depending on the host. For example, Llama 3.3 70B costs $0.59/$0.79 on Groq but $1.04/$1.04 on Together. Use the Compare page to side-by-side the same model across providers.
Deploy this open model on rented GPUs
Open weights mean you can self-host instead of paying per-token API prices. These platforms let you serve it on demand — often cheaper at scale.
Partner links — we may earn a commission if you sign up, at no cost to you. We only list platforms we'd recommend regardless.