Fireworks
Hosting provider · 15 models tracked · Founded 2022
Inference platform serving open-weight models (Llama, Qwen, DeepSeek, Kimi, GLM, MiniMax, gpt-oss) on its own infrastructure at first-party per-token pricing — distinct rates worth comparing per host.
Fireworks pricing at a glanceAugust 2026 · $ per 1M tokens
Fireworks API pricing (August 2026): 15 current models range from $0.070 to $1.74 per 1M input tokens and $0.280 to $4.40 per 1M output tokens. The cheapest paid model is GPT OSS 20B at $0.070/1M input; the priciest is DeepSeek V4 Pro at $1.74/1M input / $3.48 output. Every price links to Fireworks's official pricing page and refreshes twice daily.
Today's Fireworks prices
15 models · cheapest blended firstSorted by blended cost (cheapest first). Prices per 1M tokens, August 2026 — every price links to Fireworks's official pricing page.
| Model | Blended* | Input | Output | Cached in | Relative cost | Context | Status |
|---|---|---|---|---|---|---|---|
| GPT OSS 20B | $0.128 |
$0.070 | $0.300 | $0.035 | 128K | Current | |
| DeepSeek V4 Flash | $0.175 |
$0.140 | $0.280 | $0.028 | 1M | Current | |
| GPT OSS 120B | $0.262 |
$0.150 | $0.600 | $0.015 | 128K | Current | |
| MiniMax 2.5 | $0.525 |
$0.300 | $1.20 | $0.030 | 128K | Current | |
| MiniMax 2.7 | $0.525 |
$0.300 | $1.20 | $0.060 | 128K | Current | |
| MiniMax M3 | $0.525 |
$0.300 | $1.20 | $0.060 | 1M | Current | |
| Qwen 3.7 Plus | $0.700 |
$0.400 | $1.60 | $0.080 | 128K | Current | |
| NVIDIA Nemotron 3 Ultra | $1.05 |
$0.600 | $2.40 | $0.120 | 128K | Current | |
| Qwen 3.6 Plus | $1.13 |
$0.500 | $3.00 | $0.100 | 128K | Current | |
| Kimi K2.5 | $1.20 |
$0.600 | $3.00 | $0.100 | 256K | Current | |
| Kimi K2.6 | $1.71 |
$0.950 | $4.00 | $0.160 | 256K | Current | |
| Kimi K2.7 Code | $1.71 |
$0.950 | $4.00 | $0.190 | 256K | Current | |
| GLM 5.1 | $2.15 |
$1.40 | $4.40 | $0.260 | 128K | Current | |
| GLM 5.2 | $2.15 |
$1.40 | $4.40 | $0.140 | 1M | Current | |
| DeepSeek V4 Pro | $2.17 |
$1.74 | $3.48 | $0.145 | 1M | Current |
Quick stats
- Models tracked
- 15
- Type
- hosting
- Founded
- 2022
- Cheapest model
- GPT OSS 20B
- Cheapest blended
- $0.128/M
Open-weight models
13 open-weight models available from this provider.
Self-host Fireworks's open-weight models
These models ship with open weights, so you can serve them yourself on rented GPUs instead of paying per-token API prices — often cheaper at scale.
Partner links — we may earn a commission if you sign up, at no cost to you. We only list platforms we'd recommend regardless.