ModelPriceWatch.com
Last scan 2026-09-19 Models tracked 263 Providers 34 Cheapest paid Granite 4.0 H Micro $0.017/Mtok in Every price links to its source

GLM-5.3-FlashX

by Z.AI · 320B (A18B) parameters

Current new · 25d budget open weights cheap tier

Today's price · per 1M tokens

Input

$0.370

per 1M tokens

Output

$1.25

per 1M tokens

Blended

$0.590

blended $/1M — 3:1 weighted input:output

Cached input

$0.075

20% of input — prompt caching

How this price is scoped: Standard list pricing, per 1M tokens. Z.AI's own rate card (docs.z.ai/guides/overview/pricing) prints one value per cell for GLM-5.3-FlashX under "Latest Models" — input $0.37, cached input $0.075, output $1.25 — with no discounted/struck-through pair of the kind it showed throughout the GLM-5.3-Flash launch promotion, and Z.AI's own OpenRouter endpoint reports discount 0. Cached-input STORAGE is separately labelled "Limited-time Free" with no end date, and is not a rate we publish.

Source: official Z.AI pricing · read MODELPRICEWATCH.COM · 2026-09-19

Price receipt

We read Z.AI's own pricing page on and found GLM-5.3-FlashX listed with a price on the same row — that read is where the number above comes from, recorded as a new model. We keep the page text we read; its content hash is 63526bfdd8.

Overview

The 200 tokens/s serving tier of GLM-5.3-Flash rather than a separate model — Z.AI documents the pair on one page and publishes no standalone FlashX model card. 320B total / 18B activated, 1M-token context, 128K max output, thinking always on, and video, image, text and file input. $0.37/$1.25 per 1M tokens with cached input at $0.075 — roughly 2.5x plain GLM-5.3-Flash for the same MIT-licensed weights at huggingface.co/zai-org/GLM-5.3-Flash.

Run it yourself

Deploy this open model on rented GPUs

Open weights mean you can self-host instead of paying per-token API prices. These platforms let you serve it on demand — often cheaper at scale.

Partner links — we may earn a commission if you sign up, at no cost to you. We only list platforms we'd recommend regardless.

Capabilities

struck through = not supported
Input 3/4
Text Image Audio Video
Output
Text
Features 5/9
Prompt caching Reasoning Coding Fast inference Long context Open weights Multimodal Web search Realtime

Specifications

Provider
Z.AI
Context window
1M tokens
Modality
text, image, video
Parameters
320B (A18B)
Open source
Yes — open weights available
Released
Status
Current
Last updated
Tags
open-weightsbudgetfastmultimodalcachinglong-context

Availability verified: listed on Z.AI's own page