GLM-4-32B-0414 vs GLM-5.3-Flash
Side-by-side comparison of API pricing, specs, benchmarks, and capabilities
| Specification |
by Z.AI
|
GLM-5.3-FlashIntro price
by Z.AI
|
|---|---|---|
| Overview | ||
| Status | Current budget Open weights | Current budget Open weights |
| Released | Apr 1, 2025 | Aug 25, 2026 |
| Pricing per million tokens | ||
| Input | $0.100/Mtok | $0.075/Mtok |
| Output | $0.100/Mtok | $0.250/Mtok |
| Blended avg | $0.100/Mtok | $0.119/Mtok |
| Cached input | — | $0.015/Mtok |
| Price basis | — | Launch/intro pricing, per 1M tokens. Z.AI's own rate card (docs.z.ai/guides/overview/pricing) prints both prices for this model, with the list price struck through: "GLM-5.3-Flash is available at a 50% discount (strikethrough prices are list prices). The promotion ends at 24:00 on September 9, 2026 (UTC+8, Singapore time)." That is input $0.15 -> $0.075, cached input $0.03 -> $0.015, output $0.50 -> $0.25. The figures shown here are the current effective discounted rates — what you would pay today. On 2026-09-10 they revert to the $0.15 / $0.015 cached / $0.50 list price, and we track that reversion as a price increase with its own reversion date so this row cannot quietly go stale. Cached-input STORAGE is separately labelled "Limited-time Free" with no end date, and is not a rate we publish. Checked live on the vendor's page twice, 2026-08-26 and 2026-08-27, identical both times; Z.AI has not extended the promotion. |
| Specifications | ||
| Context window | 128K tokens | 1M tokens |
| Parameters | 32B | 320B (A18B) |
| Speed (TPS) | — | — |
| Modalities | ||
| Input |
text
|
textimage
|
| Benchmarks sources: LMSYS Chatbot Arena (UC Berkeley) / Vellum LLM Leaderboard | ||
| Independent composite scores measured for both models — each on its own scale | ||
| LMSYS Chatbot Arena | 1342.9 ELO | 1469.4 ELO |
| AA Intelligence Index | — | 57.5 |
| AA Agentic Index | — | 58.2 |
| AA-Omniscience Index (−100–100) | — | 7.5 |
| GDPval-AA v2 | — | 1763.8 ELO |
| Per-benchmark results published for GLM-5.3-Flash only — no independent per-benchmark scores exist for GLM-4-32B-0414 yet | ||
| Avg benchmark score | — | 69.8 |
| Perf / dollar | — | 587.8 |
| Terminal-Bench 2.1 | — | 84.3 |
| Humanity's Last Exam | — | 55.3 |
| AutoBench | — | 48.8 |
| Providers | ||
| Available from |
Z.AI — $0.100/$0.100/Mtok
|
Z.AI — $0.075/$0.250/Mtok
|
Cost at scale
1M tokens · 50/50 input/output| Volume | GLM-4-32B-0414 | GLM-5.3-Flash | Savings |
|---|---|---|---|
| 1M tokens | $0.1 | $0.12 | $0.02 (16.8%) |
| 10M tokens | $1 | $1.19 | $0.19 (16%) |
| 100M tokens | $10 | $11.88 | $1.88 (15.8%) |
| 1000M tokens | $100 | $118.75 | $18.75 (15.8%) |
When to pick which
distilled from the pricing and spec data above- cost dominates: $0.100/Mtok blended vs $0.119 — 15.8% less on the same 50/50 token mix
- your workload re-reads context (agents, RAG, long chats): cached input costs $0.015/Mtok — 80% off its list input price, a discount GLM-4-32B-0414 doesn't offer
- you need the longer context: 1M tokens vs 128K (7.8×)
The cheaper option here — GLM-4-32B-0414 costs $0.100/Mtok blended on Z.AI.
Summary
GLM-4-32B-0414 by Z.AI costs $0.100/Mtok input and $0.100/Mtok output, with a 128K-token context window. It supports text input.
GLM-5.3-Flash by Z.AI costs $0.075/Mtok input and $0.250/Mtok output, with a 1M-token context window. It supports text, image input.
On a blended cost basis, GLM-4-32B-0414 is 15.8% cheaper than GLM-5.3-Flash.
The two aren't directly comparable on average benchmark score: GLM-5.3-Flash has published per-benchmark results, while GLM-4-32B-0414 does not yet — it is measured today on independent composites (see the table above).
Note: Pricing is per million tokens. Actual costs vary with usage patterns, prompt caching, and batch discounts. Always verify against official provider pricing pages.