ModelPriceWatch.com
Last scan 2026-08-09 Models tracked 198 Providers 30 Cheapest paid Granite 4.0 Micro $0.017/Mtok in Every price links to its source

Groq

Hosting provider · 6 models tracked · Founded 2016

Inference platform running open models on custom LPU chips for extreme speed. Up to 1000 tokens/sec, cheapest hosted prices for many open models.

Groq pricing at a glanceAugust 2026 · $ per 1M tokens

Groq API pricing (August 2026): 6 current models range from $0.050 to $0.600 per 1M input tokens and $0.080 to $3.00 per 1M output tokens. The cheapest paid model is Llama 3.1 8B Instant at $0.050/1M input; the priciest is Qwen 3.6 27B at $0.600/1M input / $3.00 output. Every price links to Groq's official pricing page and refreshes twice daily.

Models6
Input range$0.050–$0.600
Output range$0.080–$3.00
Cached-input tiers0

Today's Groq prices

6 models · cheapest blended first

Sorted by blended cost (cheapest first). Prices per 1M tokens, August 2026 — every price links to Groq's official pricing page.

Current Groq model prices per 1M tokens, sorted by blended cost
Model Blended* Input Output Cached in Relative cost Context Status
Llama 3.1 8B Instant
$0.058
$0.050 $0.080
128K Deprecated
GPT-OSS 20B
$0.131
$0.075 $0.300
128K Current
GPT-OSS 120B
$0.262
$0.150 $0.600
128K Current
Qwen3 32B
$0.365
$0.290 $0.590
128K Current
Llama 3.3 70B Versatile
$0.640
$0.590 $0.790
128K Deprecated
Qwen 3.6 27B
$1.20
$0.600 $3.00
128K Current
* Blended = (3×input + 1×output) ÷ 4 $/1M tokens · cheap · mid · expensive MODELPRICEWATCH.COM · 2026-08-09

Quick stats

Models tracked
6
Type
hosting
Founded
2016
Cheapest model
Llama 3.1 8B Instant
Cheapest blended
$0.058/M

Open-weight models

6 open-weight models available from this provider.

GPT-OSS 120B
$0.150/M in · 128K ctx
GPT-OSS 20B
$0.075/M in · 128K ctx
Llama 3.1 8B Instant
$0.050/M in · 128K ctx
Llama 3.3 70B Versatile
$0.590/M in · 128K ctx
Qwen 3.6 27B
$0.600/M in · 128K ctx
Qwen3 32B
$0.290/M in · 128K ctx

Try Groq

Sign up and start building with Groq models.

Get started
Run it yourself

Self-host Groq's open-weight models

These models ship with open weights, so you can serve them yourself on rented GPUs instead of paying per-token API prices — often cheaper at scale.

Partner links — we may earn a commission if you sign up, at no cost to you. We only list platforms we'd recommend regardless.