ModelPriceWatch.com
Last scan 2026-08-28 Models tracked 237 Providers 31 Cheapest paid Granite 4.0 H Micro $0.017/Mtok in Every price links to its source

GLM-4-32B-0414 vs GLM-5.3-Flash

Side-by-side comparison of API pricing, specs, benchmarks, and capabilities

GLM-4-32B-0414 is 15.8% cheaper on blended cost ($0.100 vs $0.119/Mtok)
Specification
by Z.AI
GLM-5.3-FlashIntro price
by Z.AI
Overview
StatusCurrent budget Open weights Current budget Open weights
Released Apr 1, 2025 Aug 25, 2026
Pricing per million tokens
Input $0.100/Mtok $0.075/Mtok
Output $0.100/Mtok $0.250/Mtok
Blended avg $0.100/Mtok $0.119/Mtok
Cached input $0.015/Mtok
Price basis Launch/intro pricing, per 1M tokens. Z.AI's own rate card (docs.z.ai/guides/overview/pricing) prints both prices for this model, with the list price struck through: "GLM-5.3-Flash is available at a 50% discount (strikethrough prices are list prices). The promotion ends at 24:00 on September 9, 2026 (UTC+8, Singapore time)." That is input $0.15 -> $0.075, cached input $0.03 -> $0.015, output $0.50 -> $0.25. The figures shown here are the current effective discounted rates — what you would pay today. On 2026-09-10 they revert to the $0.15 / $0.015 cached / $0.50 list price, and we track that reversion as a price increase with its own reversion date so this row cannot quietly go stale. Cached-input STORAGE is separately labelled "Limited-time Free" with no end date, and is not a rate we publish. Checked live on the vendor's page twice, 2026-08-26 and 2026-08-27, identical both times; Z.AI has not extended the promotion.
Specifications
Context window 128K tokens 1M tokens
Parameters 32B 320B (A18B)
Speed (TPS)
Modalities
Input
text
textimage
Benchmarks sources: LMSYS Chatbot Arena (UC Berkeley) / Vellum LLM Leaderboard
Independent composite scores measured for both models — each on its own scale
LMSYS Chatbot Arena 1342.9 ELO 1469.4 ELO
AA Intelligence Index 57.5
AA Agentic Index 58.2
AA-Omniscience Index (−100–100) 7.5
GDPval-AA v2 1763.8 ELO
Per-benchmark results published for GLM-5.3-Flash only — no independent per-benchmark scores exist for GLM-4-32B-0414 yet
Avg benchmark score 69.8
Perf / dollar 587.8
Terminal-Bench 2.1 84.3
Humanity's Last Exam 55.3
AutoBench 48.8
Providers
Available from
Z.AI — $0.100/$0.100/Mtok
Z.AI — $0.075/$0.250/Mtok

Cost at scale

1M tokens · 50/50 input/output
Projected cost of GLM-4-32B-0414 vs GLM-5.3-Flash at increasing token volumes
VolumeGLM-4-32B-0414GLM-5.3-FlashSavings
1M tokens $0.1 $0.12 $0.02 (16.8%)
10M tokens $1 $1.19 $0.19 (16%)
100M tokens $10 $11.88 $1.88 (15.8%)
1000M tokens $100 $118.75 $18.75 (15.8%)

When to pick which

distilled from the pricing and spec data above
Pick GLM-4-32B-0414 if…
  • cost dominates: $0.100/Mtok blended vs $0.119 — 15.8% less on the same 50/50 token mix
Pick GLM-5.3-Flash if…
  • your workload re-reads context (agents, RAG, long chats): cached input costs $0.015/Mtok — 80% off its list input price, a discount GLM-4-32B-0414 doesn't offer
  • you need the longer context: 1M tokens vs 128K (7.8×)
Try GLM-4-32B-0414 on Z.AI

The cheaper option here — GLM-4-32B-0414 costs $0.100/Mtok blended on Z.AI.

Get API key →

Summary

GLM-4-32B-0414 by Z.AI costs $0.100/Mtok input and $0.100/Mtok output, with a 128K-token context window. It supports text input.

GLM-5.3-Flash by Z.AI costs $0.075/Mtok input and $0.250/Mtok output, with a 1M-token context window. It supports text, image input.

On a blended cost basis, GLM-4-32B-0414 is 15.8% cheaper than GLM-5.3-Flash.

The two aren't directly comparable on average benchmark score: GLM-5.3-Flash has published per-benchmark results, while GLM-4-32B-0414 does not yet — it is measured today on independent composites (see the table above).

Note: Pricing is per million tokens. Actual costs vary with usage patterns, prompt caching, and batch discounts. Always verify against official provider pricing pages.

More comparisons

pairs sharing a model with this page