ModelPriceWatch.com
Last scan 2026-09-22 Models tracked 267 Providers 35 Cheapest paid Granite 4.0 H Micro $0.017/Mtok in Every price links to its source

Open-weight models

83 open-weight models hosted across inference providers. Compare hosting prices — the model weights are free, you pay only for compute.

Open-weight models

83

current, verified hosted prices

Cheapest input /Mtok

$0

GLM-4.5-Flash

Fastest inference

1000 TPS

GPT-OSS 20B

Largest context

10M tokens

Llama 4 Scout

Today's open-weight prices

83 models · sorted by input price, cheapest first
Open-weight LLM hosting prices per million tokens, sorted by input price — the same model can be listed by several hosts
Parameters Notes
GLM-4.5-FlashZ.AI $0 $0 128K
GLM-4.6V-FlashZ.AI $0 $0 128K
GLM-4.7-FlashZ.AI $0 $0 128K
Granite 4.0 H MicroIBM $0.017 $0.112 128K
LFM2.5 8B A1BTogether $0.030 $0.120 33K 8.5B (A1B)
GLM-OCRZ.AI $0.030 $0.030 128K
GLM-4.6V-FlashXZ.AI $0.040cached $0.004 $0.400 128K
Hy-MT2 1.8BTencent $0.044 $0.177 8K 1.8B
Llama 3.1 8BMeta $0.050 $0.080 128K 8B
Granite 4 H SmallIBM $0.064 $0.265 128K
Baichuan M2-32BBaichuan $0.070 $0.070 131K 32B
GPT-OSS 20BFireworks $0.070cached $0.035 $0.300 128K 20B
GLM-4.7-FlashXZ.AI $0.070cached $0.010 $0.400 128K
Hy-MT2 30B-A3BTencent $0.074 $0.295 8K 30B (A3B)
GPT-OSS 20BGroq $0.075 $0.300 131K 20B 1000 TPS
Qwen3-32BDeepInfra $0.080 $0.280 128K 32B
NVIDIA Nemotron 3.5 LightningDeepInfra $0.080cached $0.040 $0.200 256K 30B (3B active)
Ministral 3 3BMistral $0.100 $0.100 128K 3B
Voxtral Small 24BMistral $0.100 $0.400 128K 24B
GLM-4-32B-0414Z.AI $0.100 $0.100 128K 32B
Llama 4 ScoutDeepInfra $0.100 $0.300 10M 17B (16 experts)
Granite Embedding 278M MultilingualIBM $0.106 278M
Llama 4 ScoutMeta $0.110 $0.340 10M 17B (16 experts)
GPT-OSS 120BFireworks $0.150cached $0.015 $0.600 128K 120B
GPT-OSS 120BGroq $0.150 $0.600 131K 120B 500 TPS
GPT-OSS 120BTogether $0.150 $0.600 131K 120B
GLM-5.3-FlashZ.AI $0.150cached $0.030 $0.500 1M 320B (A18B)
Qwen3-32BAlibaba $0.160 $0.640 131K 32B
Jamba MiniAI21 Labs $0.200 $0.400 256K
GLM-4.5-AirZ.AI $0.200cached $0.030 $1.10 128K
Llama 4 MaverickDeepInfra $0.200 $0.800 1M 17B (128 experts)
Qwen3.5-35B-A3BAlibaba $0.250 $2.00 262K 35B (A3B)
Qwen3-32BGroq $0.290 $0.590 128K 32B 662 TPS
Qwen3.5-27BAlibaba $0.300 $2.40 262K 27B
MiniMax-M2.7Fireworks $0.300cached $0.060 $1.20 128K
MiniMax-M3Fireworks $0.300cached $0.060 $1.20 1M
MiniMax-M2.7MiniMax $0.300cached $0.060 $1.20 205K
MiniMax-M3MiniMax $0.300cached $0.060 $1.20 1M
MiniMax-M3Together $0.300cached $0.060 $1.20 524K
GLM-4.6VZ.AI $0.300cached $0.050 $0.900 128K
Inkling SmallThinking Machines $0.300cached $0.060 $1.20 262K 276B total / 12B active
Qwen3.7-PlusTogether $0.320 $1.28 1M
Muse Glimmer 30BFireworks $0.350cached $0.040 $1.50 131K 30B
Muse Glimmer 30BTogether $0.350cached $0.040 $1.50 131K 30B
GLM-5.3-FlashXZ.AI $0.370cached $0.075 $1.25 1M 320B (A18B)
Qwen3.5-122B-A10BAlibaba $0.400 $3.20 262K 122B (A10B)
Qwen3.7-PlusFireworks $0.400cached $0.080 $1.60 128K
Qwen3.8-27BAlibaba $0.500 $3.00 1M 27B
Llama 3.3 70BMeta $0.590 $0.790 128K 70B
Qwen3.6-27BAlibaba $0.600 $3.60 262K 27B
Qwen3.5-397B-A17BAlibaba $0.600 $3.60 262K 397B (A17B)
NVIDIA Nemotron 3 UltraFireworks $0.600cached $0.120 $2.40 128K
Qwen3.6-27BGroq $0.600 $3.00 131K 27B 500 TPS
Kimi K2.5Moonshot $0.600cached $0.100 $3.00 262K
GLM-4.5Z.AI $0.600cached $0.110 $2.20 128K
GLM-4.6Z.AI $0.600cached $0.110 $2.20 200K
GLM-4.5VZ.AI $0.600cached $0.110 $1.80 106B (A12B)
GLM-4.7Z.AI $0.600cached $0.110 $2.20 128K
LongCat-2.0Meituan $0.750cached $0.015 $2.95 262K
Hy4 previewTencent $0.834cached $0.042 $2.50 1M 770B (A49B)
Kimi K2.6Fireworks $0.950cached $0.160 $4.00 256K
Kimi K2.7 CodeFireworks $0.950cached $0.190 $4.00 256K
Kimi K2.6Moonshot $0.950cached $0.160 $4.00 262K
Kimi K2.7 CodeMoonshot $0.950cached $0.190 $4.00 262K
GLM-5Z.AI $1.00cached $0.200 $3.20 128K
InklingThinking Machines $1.00cached $0.170 $4.05 262K 975B total / 41B active
Llama 3.3 70BTogether $1.04 $1.04 131K 70B
GLM-4.5-AirXZ.AI $1.10cached $0.220 $4.50 128K
GLM-5-TurboZ.AI $1.20cached $0.240 $4.00 128K
GLM-5V-TurboZ.AI $1.20cached $0.240 $4.00 128K
GLM-5.1Fireworks $1.40cached $0.260 $4.40 128K
GLM-5.2Fireworks $1.40cached $0.140 $4.40 1M
GLM-5.3Fireworks $1.40cached $0.260 $4.40 1M
GLM-5.2Together $1.40cached $0.260 $4.40 1M
GLM-5.3Together $1.40cached $0.260 $4.40 1M
GLM-5.1Z.AI $1.40cached $0.260 $4.40 128K
GLM-5.2Z.AI $1.40cached $0.260 $4.40 1M
GLM-5.3Z.AI $1.40cached $0.260 $4.40 1M
Kimi K2.7 Code HighSpeedMoonshot $1.90cached $0.380 $8.00 262K
Jamba LargeAI21 Labs $2.00 $8.00 256K
Qwen3.8-2.4T-A95BAlibaba $2.00 $6.00 1M 2.4T (A95B)
GLM-4.5-XZ.AI $2.20cached $0.450 $8.90 128K
Kimi K3Moonshot $3.00cached $0.300 $15.00 1M
Prices per 1M tokens in USD · open-weight models list the cheapest verified hosted price · how we verify

MODELPRICEWATCH.COM · 2026-09-22

How open-weight pricing works

Open weight models have freely available model weights — anyone can download and run them. You pay only for the compute to serve inference. Hosting providers like Groq (custom LPU chips, fastest inference), Together AI (200+ models, competitive pricing), and Fireworks offer per-token pricing without lock-in.

The same model can have very different prices depending on the host. For example, Llama 3.3 70B costs $0.59/$0.79 on Groq but $1.04/$1.04 on Together.

Run it yourself

Deploy this open model on rented GPUs

Open weights mean you can self-host instead of paying per-token API prices. These platforms let you serve it on demand — often cheaper at scale.

Partner links — we may earn a commission if you sign up, at no cost to you. We only list platforms we'd recommend regardless.