ModelPriceWatch.com
Last scan 2026-08-09 Models tracked 198 Providers 30 Cheapest paid Granite 4.0 Micro $0.017/Mtok in Every price links to its source

Open-weight models

70 open-weight models hosted across inference providers. Compare hosting prices — the model weights are free, you pay only for compute.

Open-weight models

70

current, verified hosted prices

Cheapest input /Mtok

$0

GLM-4.7-Flash

Fastest inference

1000 TPS

GPT-OSS 20B

Largest context

10M tokens

Llama 4 Scout

Today's open-weight prices

70 models · sorted by input price, cheapest first
Open-weight LLM hosting prices per million tokens, sorted by input price — the same model can be listed by several hosts
Parameters Notes
GLM-4.7-FlashZ.AI $0 $0 128K
Granite 4.0 MicroIBM $0.017 $0.112 128K
GLM-OCRZ.AI $0.030 $0.030 128K
Llama 3.1 8B InstantGroq $0.050 $0.080 128K 8B 840 TPS
Llama 3.1 8BMeta $0.050 $0.080 128K 8B
Granite 4 H SmallIBM $0.060 $0.250 128K
Baichuan M2-32BBaichuan $0.070 $0.070 33K 32B
GPT OSS 20BFireworks $0.070cached $0.035 $0.300 128K 20B
GLM-4.7-FlashXZ.AI $0.070cached $0.010 $0.400 128K
GPT-OSS 20BGroq $0.075 $0.300 128K 20B 1000 TPS
Qwen3 32BDeepInfra $0.080 $0.280 128K 32B
Ministral 3 3BMistral $0.100 $0.100 128K 3B
Voxtral Small 24BMistral $0.100 $0.300 128K 24B
GLM-4-32B-0414Z.AI $0.100 $0.100 128K 32B
Llama 4 ScoutDeepInfra $0.100 $0.300 10M 17B (16 experts)
Granite Embedding 278M MultilingualIBM $0.106 $— 278M
Llama 4 ScoutMeta $0.110 $0.340 10M 17B (16 experts)
GPT OSS 120BFireworks $0.150cached $0.015 $0.600 128K 120B
GPT-OSS 120BGroq $0.150 $0.600 128K 120B 500 TPS
Granite 4 H MediumIBM $0.150 $0.600 128K
gpt-oss-120BTogether $0.150 $0.600 128K 120B
Qwen3-32BAlibaba $0.160 $0.640 131K 32B
Jamba MiniAI21 Labs $0.200 $0.400 256K
GLM-4.5-AirZ.AI $0.200cached $0.030 $1.10 128K
Llama 4 MaverickDeepInfra $0.200 $0.800 1M 17B (128 experts)
Gemma-4-31B-it-PearlTogether $0.280 $0.860 128K 31B
Qwen3 32BGroq $0.290 $0.590 128K 32B 662 TPS
MiniMax 2.5Fireworks $0.300cached $0.030 $1.20 128K
MiniMax 2.7Fireworks $0.300cached $0.060 $1.20 128K
MiniMax M3Fireworks $0.300cached $0.060 $1.20 1M
Granite 4 H LargeIBM $0.300 $1.20 128K
MiniMax-M2.7MiniMax $0.300cached $0.060 $1.20 205K
MiniMax-M3MiniMax $0.300cached $0.060 $1.20 1M
MiniMax M3Together $0.300cached $0.060 $1.20 1M
GLM-4.6VZ.AI $0.300cached $0.050 $0.900 128K
Inkling SmallThinking Machines $0.300cached $0.060 $1.20 262K 276B total / 12B active
Qwen3.7-PlusTogether $0.320 $1.28 128K
Qwen 3.7 PlusFireworks $0.400cached $0.080 $1.60 128K
Qwen 3.6 PlusFireworks $0.500cached $0.100 $3.00 128K
Llama 3.3 70B VersatileGroq $0.590 $0.790 128K 70B 394 TPS
Llama 3.3 70BMeta $0.590 $0.790 128K 70B
Qwen3.6-27BAlibaba $0.600 $3.60 262K 27B
Qwen3.5-397B-A17BAlibaba $0.600 $3.60 262K 397B (A17B)
Kimi K2.5Fireworks $0.600cached $0.100 $3.00 256K
NVIDIA Nemotron 3 UltraFireworks $0.600cached $0.120 $2.40 128K
Qwen 3.6 27BGroq $0.600 $3.00 128K 27B 500 TPS
Kimi K2.5Moonshot $0.600cached $0.100 $3.00 262K
NVIDIA Nemotron 3 UltraTogether $0.600cached $0.200 $3.60 128K
GLM-4.5Z.AI $0.600cached $0.110 $2.20 128K
GLM-4.6Z.AI $0.600cached $0.110 $2.20 200K
GLM-4.7Z.AI $0.600cached $0.110 $2.20 128K
LongCat-2.0Meituan $0.750cached $0.015 $2.95 262K
Kimi K2.6Fireworks $0.950cached $0.160 $4.00 256K
Kimi K2.7 CodeFireworks $0.950cached $0.190 $4.00 256K
Kimi K2.6Moonshot $0.950cached $0.160 $4.00 262K
Kimi K2.7 CodeMoonshot $0.950cached $0.190 $4.00 262K
Kimi K2.7 CodeTogether $0.950cached $0.190 $4.00 256K
GLM-5Z.AI $1.00cached $0.200 $3.20 128K
InklingThinking Machines $1.00cached $0.170 $4.05 262K 975B total / 41B active
Llama 3.3 70BTogether $1.04 $1.04 128K 70B
GLM-5-TurboZ.AI $1.20cached $0.240 $4.00 128K
GLM-5V-TurboZ.AI $1.20cached $0.240 $4.00 128K
GLM 5.1Fireworks $1.40cached $0.260 $4.40 128K
GLM 5.2Fireworks $1.40cached $0.140 $4.40 1M
GLM-5.2Together $1.40cached $0.260 $4.40 1M
GLM-5.1Z.AI $1.40cached $0.260 $4.40 128K
GLM-5.2Z.AI $1.40cached $0.260 $4.40 1M
Kimi K2.7 Code HighSpeedMoonshot $1.90cached $0.380 $8.00 262K
Jamba LargeAI21 Labs $2.00 $8.00 256K
Kimi K3Moonshot $3.00cached $0.300 $15.00 1M
Prices per 1M tokens in USD · open-weight models list the cheapest verified hosted price · how we verify

MODELPRICEWATCH.COM · 2026-08-09

How open-weight pricing works

Open weight models have freely available model weights — anyone can download and run them. You pay only for the compute to serve inference. Hosting providers like Groq (custom LPU chips, fastest inference), Together AI (200+ models, competitive pricing), and Fireworks offer per-token pricing without lock-in.

The same model can have very different prices depending on the host. For example, Llama 3.3 70B costs $0.59/$0.79 on Groq but $1.04/$1.04 on Together. Use the Compare page to side-by-side the same model across providers.

Run it yourself

Deploy this open model on rented GPUs

Open weights mean you can self-host instead of paying per-token API prices. These platforms let you serve it on demand — often cheaper at scale.

Partner links — we may earn a commission if you sign up, at no cost to you. We only list platforms we'd recommend regardless.