Groq
Hosting provider · 6 models tracked · Founded 2016
Inference platform running open models on custom LPU chips for extreme speed. Up to 1000 tokens/sec, cheapest hosted prices for many open models.
Groq pricing at a glanceAugust 2026 · $ per 1M tokens
Groq API pricing (August 2026): 6 current models range from $0.050 to $0.600 per 1M input tokens and $0.080 to $3.00 per 1M output tokens. The cheapest paid model is Llama 3.1 8B Instant at $0.050/1M input; the priciest is Qwen 3.6 27B at $0.600/1M input / $3.00 output. Every price links to Groq's official pricing page and refreshes twice daily.
Today's Groq prices
6 models · cheapest blended firstSorted by blended cost (cheapest first). Prices per 1M tokens, August 2026 — every price links to Groq's official pricing page.
| Model | Blended* | Input | Output | Cached in | Relative cost | Context | Status |
|---|---|---|---|---|---|---|---|
| Llama 3.1 8B Instant | $0.058 |
$0.050 | $0.080 | — | 128K | Deprecated | |
| GPT-OSS 20B | $0.131 |
$0.075 | $0.300 | — | 128K | Current | |
| GPT-OSS 120B | $0.262 |
$0.150 | $0.600 | — | 128K | Current | |
| Qwen3 32B | $0.365 |
$0.290 | $0.590 | — | 128K | Current | |
| Llama 3.3 70B Versatile | $0.640 |
$0.590 | $0.790 | — | 128K | Deprecated | |
| Qwen 3.6 27B | $1.20 |
$0.600 | $3.00 | — | 128K | Current |
Quick stats
- Models tracked
- 6
- Type
- hosting
- Founded
- 2016
- Cheapest model
- Llama 3.1 8B Instant
- Cheapest blended
- $0.058/M
Open-weight models
6 open-weight models available from this provider.
Self-host Groq's open-weight models
These models ship with open weights, so you can serve them yourself on rented GPUs instead of paying per-token API prices — often cheaper at scale.
Partner links — we may earn a commission if you sign up, at no cost to you. We only list platforms we'd recommend regardless.