Today's price · per 1M tokens
Input
$1.40
per 1M tokens
Output
$4.40
per 1M tokens
Blended
$2.90
avg of input & output $/1M
Cached input
$0.260
19% of input — prompt caching
Available on 2 hosts
cheapest blended first · $ per 1M tokensGLM 5.1 is sold by 2 providers. Prices are per 1M tokens (blended = avg of input & output). The first-party row is the model maker; “vs first-party” shows each host’s blended price relative to it.
| Host | Input | Output | Blended | vs first-party |
|---|---|---|---|---|
| Fireworks details → | $1.40 | $4.40 | $2.90 | 0% |
| Z.AI first-party | $1.40 | $4.40 | $2.90 | — |
Overview
GLM-5.1 model. $1.40/$4.40 per 1M; cached input $0.26.
Deploy this open model on rented GPUs
Open weights mean you can self-host instead of paying per-token API prices. These platforms let you serve it on demand — often cheaper at scale.
Partner links — we may earn a commission if you sign up, at no cost to you. We only list platforms we'd recommend regardless.
Capabilities
struck through = not supportedBenchmark performance
accuracy % · higher is betterEvery per-benchmark score we hold for this model is reported by its own vendor, so none is counted toward an accuracy average or any ranking. The percentile above is its standing across the independent composites below.
- LMSYS Chatbot Arena: 1468.8 ELO (Human preference)
- Ranks #20 of 133 comparably-measured models by percentile score, across 1 independent measurement
- Ranks #54 of 133 comparably-tested models by normalized performance per dollar
160 tokens/sec output
Source: Z.AI model card · updated Mar 1, 2026 · See full rankings →
Specifications
- Provider
- Z.AI
- Context window
- 128K tokens
- Modality
- text
- Parameters
- Proprietary
- Open source
- Yes — open weights available
- Released
- Oct 1, 2025
- Status
- Current
- Last updated
- Jun 25, 2026
- Tags
Availability verified: Jul 20, 2026 — listed on Z.AI's own page