Today's price · per 1M tokens
Input
$0.300
per 1M tokens
Output
$1.20
per 1M tokens
Blended
$0.525
blended $/1M — 3:1 weighted input:output
Cached input
$0.060
20% of input — prompt caching
Price receipt
We read MiniMax's own pricing page on and found MiniMax-M3 listed with a price on the same row — that read is where the number above comes from. We keep the page text we read; its content hash is 5ccab58c6b.
Available on 3 hosts
cheapest blended first · $ per 1M tokensMiniMax-M3 is sold by 3 providers. Prices are per 1M tokens (blended = (3×input + 1×output) ÷ 4). The first-party row is the model maker; “vs first-party” shows each host’s blended price relative to it.
| Host | Input | Output | Blended | vs first-party |
|---|---|---|---|---|
| Fireworks details → | $0.300 | $1.20 | $0.525 | 0% |
| MiniMax first-party | $0.300 | $1.20 | $0.525 | — |
| Together details → | $0.300 | $1.20 | $0.525 | 0% |
Overview
MiniMax's open-weight flagship built on its MSA sparse-attention architecture, combining frontier-level coding, a 1M-token context, and native multimodality. $0.30/$1.20 per 1M tokens.
Deploy this open model on rented GPUs
Open weights mean you can self-host instead of paying per-token API prices. These platforms let you serve it on demand — often cheaper at scale.
Partner links — we may earn a commission if you sign up, at no cost to you. We only list platforms we'd recommend regardless.
Capabilities
struck through = not supportedBenchmark performance
accuracy % · higher is better- AA Intelligence Index: 29.6 (Artificial Analysis composite across reasoning, knowledge and coding evals)
- AA-Omniscience Index: 1.4 (−100–100)
- GDPval-AA v2: 1304 ELO (Real-world work tasks, human baseline = 1000)
- LMSYS Chatbot Arena: 1441.3 ELO (Human preference)
- Ranks #47 of 84 benchmarked models by average score
- Ranks #89 of 181 comparably-measured models by percentile score, across 17 independent measurements
- Ranks #30 of 181 comparably-tested models by normalized performance per dollar
- Strongest at GPQA Diamond — 93%, #12 of 65
98.6 tokens/sec output 0.85s latency to first token (TTFT)
Source: Vellum LLM Leaderboard · updated · See full rankings →
Specifications
- Provider
- MiniMax
- Context window
- 1M tokens
- Modality
- text, image, video
- Parameters
- Proprietary
- Open source
- Yes — open weights available
- Released
- Status
- Current
- Last updated
- Tags
Availability verified: — listed on MiniMax's own page