Open-weight models
83 open-weight models hosted across inference providers. Compare hosting prices — the model weights are free, you pay only for compute.
Open-weight models
83
current, verified hosted prices
Cheapest input /Mtok
$0
GLM-4.5-Flash
Fastest inference
1000 TPS
GPT-OSS 20B
Largest context
10M tokens
Llama 4 Scout
Today's open-weight prices
83 models · sorted by input price, cheapest first| Parameters | Notes | ||||
|---|---|---|---|---|---|
| GLM-4.5-FlashZ.AI | $0 | $0 | 128K | — | |
| GLM-4.6V-FlashZ.AI | $0 | $0 | 128K | — | |
| GLM-4.7-FlashZ.AI | $0 | $0 | 128K | — | |
| Granite 4.0 H MicroIBM | $0.017 | $0.112 | 128K | — | |
| LFM2.5 8B A1BTogether | $0.030 | $0.120 | 33K | 8.5B (A1B) | |
| GLM-OCRZ.AI | $0.030 | $0.030 | 128K | — | |
| GLM-4.6V-FlashXZ.AI | $0.040cached $0.004 | $0.400 | 128K | — | |
| Hy-MT2 1.8BTencent | $0.044 | $0.177 | 8K | 1.8B | |
| Llama 3.1 8BMeta | $0.050 | $0.080 | 128K | 8B | |
| Granite 4 H SmallIBM | $0.064 | $0.265 | 128K | — | |
| Baichuan M2-32BBaichuan | $0.070 | $0.070 | 131K | 32B | |
| GPT-OSS 20BFireworks | $0.070cached $0.035 | $0.300 | 128K | 20B | |
| GLM-4.7-FlashXZ.AI | $0.070cached $0.010 | $0.400 | 128K | — | |
| Hy-MT2 30B-A3BTencent | $0.074 | $0.295 | 8K | 30B (A3B) | |
| GPT-OSS 20BGroq | $0.075 | $0.300 | 131K | 20B | 1000 TPS |
| Qwen3-32BDeepInfra | $0.080 | $0.280 | 128K | 32B | |
| NVIDIA Nemotron 3.5 LightningDeepInfra | $0.080cached $0.040 | $0.200 | 256K | 30B (3B active) | |
| Ministral 3 3BMistral | $0.100 | $0.100 | 128K | 3B | |
| Voxtral Small 24BMistral | $0.100 | $0.400 | 128K | 24B | |
| GLM-4-32B-0414Z.AI | $0.100 | $0.100 | 128K | 32B | |
| Llama 4 ScoutDeepInfra | $0.100 | $0.300 | 10M | 17B (16 experts) | |
| Granite Embedding 278M MultilingualIBM | $0.106 | — | — | 278M | |
| Llama 4 ScoutMeta | $0.110 | $0.340 | 10M | 17B (16 experts) | |
| GPT-OSS 120BFireworks | $0.150cached $0.015 | $0.600 | 128K | 120B | |
| GPT-OSS 120BGroq | $0.150 | $0.600 | 131K | 120B | 500 TPS |
| GPT-OSS 120BTogether | $0.150 | $0.600 | 131K | 120B | |
| GLM-5.3-FlashZ.AI | $0.150cached $0.030 | $0.500 | 1M | 320B (A18B) | |
| Qwen3-32BAlibaba | $0.160 | $0.640 | 131K | 32B | |
| Jamba MiniAI21 Labs | $0.200 | $0.400 | 256K | — | |
| GLM-4.5-AirZ.AI | $0.200cached $0.030 | $1.10 | 128K | — | |
| Llama 4 MaverickDeepInfra | $0.200 | $0.800 | 1M | 17B (128 experts) | |
| Qwen3.5-35B-A3BAlibaba | $0.250 | $2.00 | 262K | 35B (A3B) | |
| Qwen3-32BGroq | $0.290 | $0.590 | 128K | 32B | 662 TPS |
| Qwen3.5-27BAlibaba | $0.300 | $2.40 | 262K | 27B | |
| MiniMax-M2.7Fireworks | $0.300cached $0.060 | $1.20 | 128K | — | |
| MiniMax-M3Fireworks | $0.300cached $0.060 | $1.20 | 1M | — | |
| MiniMax-M2.7MiniMax | $0.300cached $0.060 | $1.20 | 205K | — | |
| MiniMax-M3MiniMax | $0.300cached $0.060 | $1.20 | 1M | — | |
| MiniMax-M3Together | $0.300cached $0.060 | $1.20 | 524K | — | |
| GLM-4.6VZ.AI | $0.300cached $0.050 | $0.900 | 128K | — | |
| Inkling SmallThinking Machines | $0.300cached $0.060 | $1.20 | 262K | 276B total / 12B active | |
| Qwen3.7-PlusTogether | $0.320 | $1.28 | 1M | — | |
| Muse Glimmer 30BFireworks | $0.350cached $0.040 | $1.50 | 131K | 30B | |
| Muse Glimmer 30BTogether | $0.350cached $0.040 | $1.50 | 131K | 30B | |
| GLM-5.3-FlashXZ.AI | $0.370cached $0.075 | $1.25 | 1M | 320B (A18B) | |
| Qwen3.5-122B-A10BAlibaba | $0.400 | $3.20 | 262K | 122B (A10B) | |
| Qwen3.7-PlusFireworks | $0.400cached $0.080 | $1.60 | 128K | — | |
| Qwen3.8-27BAlibaba | $0.500 | $3.00 | 1M | 27B | |
| Llama 3.3 70BMeta | $0.590 | $0.790 | 128K | 70B | |
| Qwen3.6-27BAlibaba | $0.600 | $3.60 | 262K | 27B | |
| Qwen3.5-397B-A17BAlibaba | $0.600 | $3.60 | 262K | 397B (A17B) | |
| NVIDIA Nemotron 3 UltraFireworks | $0.600cached $0.120 | $2.40 | 128K | — | |
| Qwen3.6-27BGroq | $0.600 | $3.00 | 131K | 27B | 500 TPS |
| Kimi K2.5Moonshot | $0.600cached $0.100 | $3.00 | 262K | — | |
| GLM-4.5Z.AI | $0.600cached $0.110 | $2.20 | 128K | — | |
| GLM-4.6Z.AI | $0.600cached $0.110 | $2.20 | 200K | — | |
| GLM-4.5VZ.AI | $0.600cached $0.110 | $1.80 | — | 106B (A12B) | |
| GLM-4.7Z.AI | $0.600cached $0.110 | $2.20 | 128K | — | |
| LongCat-2.0Meituan | $0.750cached $0.015 | $2.95 | 262K | — | |
| Hy4 previewTencent | $0.834cached $0.042 | $2.50 | 1M | 770B (A49B) | |
| Kimi K2.6Fireworks | $0.950cached $0.160 | $4.00 | 256K | — | |
| Kimi K2.7 CodeFireworks | $0.950cached $0.190 | $4.00 | 256K | — | |
| Kimi K2.6Moonshot | $0.950cached $0.160 | $4.00 | 262K | — | |
| Kimi K2.7 CodeMoonshot | $0.950cached $0.190 | $4.00 | 262K | — | |
| GLM-5Z.AI | $1.00cached $0.200 | $3.20 | 128K | — | |
| InklingThinking Machines | $1.00cached $0.170 | $4.05 | 262K | 975B total / 41B active | |
| Llama 3.3 70BTogether | $1.04 | $1.04 | 131K | 70B | |
| GLM-4.5-AirXZ.AI | $1.10cached $0.220 | $4.50 | 128K | — | |
| GLM-5-TurboZ.AI | $1.20cached $0.240 | $4.00 | 128K | — | |
| GLM-5V-TurboZ.AI | $1.20cached $0.240 | $4.00 | 128K | — | |
| GLM-5.1Fireworks | $1.40cached $0.260 | $4.40 | 128K | — | |
| GLM-5.2Fireworks | $1.40cached $0.140 | $4.40 | 1M | — | |
| GLM-5.3Fireworks | $1.40cached $0.260 | $4.40 | 1M | — | |
| GLM-5.2Together | $1.40cached $0.260 | $4.40 | 1M | — | |
| GLM-5.3Together | $1.40cached $0.260 | $4.40 | 1M | — | |
| GLM-5.1Z.AI | $1.40cached $0.260 | $4.40 | 128K | — | |
| GLM-5.2Z.AI | $1.40cached $0.260 | $4.40 | 1M | — | |
| GLM-5.3Z.AI | $1.40cached $0.260 | $4.40 | 1M | — | |
| Kimi K2.7 Code HighSpeedMoonshot | $1.90cached $0.380 | $8.00 | 262K | — | |
| Jamba LargeAI21 Labs | $2.00 | $8.00 | 256K | — | |
| Qwen3.8-2.4T-A95BAlibaba | $2.00 | $6.00 | 1M | 2.4T (A95B) | |
| GLM-4.5-XZ.AI | $2.20cached $0.450 | $8.90 | 128K | — | |
| Kimi K3Moonshot | $3.00cached $0.300 | $15.00 | 1M | — |
MODELPRICEWATCH.COM · 2026-09-22
How open-weight pricing works
Open weight models have freely available model weights — anyone can download and run them. You pay only for the compute to serve inference. Hosting providers like Groq (custom LPU chips, fastest inference), Together AI (200+ models, competitive pricing), and Fireworks offer per-token pricing without lock-in.
The same model can have very different prices depending on the host. For example, Llama 3.3 70B costs $0.59/$0.79 on Groq but $1.04/$1.04 on Together.
Deploy this open model on rented GPUs
Open weights mean you can self-host instead of paying per-token API prices. These platforms let you serve it on demand — often cheaper at scale.
Partner links — we may earn a commission if you sign up, at no cost to you. We only list platforms we'd recommend regardless.