ModelPriceWatch.com
Last scan 2026-08-28 Models tracked 237 Providers 31 Cheapest paid Granite 4.0 H Micro $0.017/Mtok in Every price links to its source

GLM-5.3-Flash vs Qwen3.8-Flash

Side-by-side comparison of API pricing, specs, benchmarks, and capabilities

GLM-5.3-Flash is 48.4% cheaper on blended cost ($0.119 vs $0.230/Mtok)
Specification
GLM-5.3-FlashIntro price
by Z.AI
Overview
StatusCurrent budget Open weights Current budget
Released Aug 25, 2026 Aug 26, 2026
Pricing per million tokens
Input $0.075/Mtok $0.150/Mtok
Output $0.250/Mtok $0.470/Mtok
Blended avg $0.119/Mtok $0.230/Mtok
Cached input $0.015/Mtok
Price basis Launch/intro pricing, per 1M tokens. Z.AI's own rate card (docs.z.ai/guides/overview/pricing) prints both prices for this model, with the list price struck through: "GLM-5.3-Flash is available at a 50% discount (strikethrough prices are list prices). The promotion ends at 24:00 on September 9, 2026 (UTC+8, Singapore time)." That is input $0.15 -> $0.075, cached input $0.03 -> $0.015, output $0.50 -> $0.25. The figures shown here are the current effective discounted rates — what you would pay today. On 2026-09-10 they revert to the $0.15 / $0.015 cached / $0.50 list price, and we track that reversion as a price increase with its own reversion date so this row cannot quietly go stale. Cached-input STORAGE is separately labelled "Limited-time Free" with no end date, and is not a rate we publish. Checked live on the vendor's page twice, 2026-08-26 and 2026-08-27, identical both times; Z.AI has not extended the promotion. Alibaba Model Studio International list price, the SKU's single 0<Token≤1M tier — unlike Qwen3.6-Flash and Qwen3.7-Flash this row is not tiered by request size. The row is labelled context-cache eligible but the table prints no cache rate for it, so cached input is left unset rather than assumed; unlike its Qwen3.7 predecessor it carries no 50% batch-inference label. A free quota of 1M tokens (90 days) applies in Singapore only and is a trial credit, not a list price.
Specifications
Context window 1M tokens 1M tokens
Parameters 320B (A18B) Proprietary
Speed (TPS)
Modalities
Input
textimage
text
Benchmarks sources: Vellum LLM Leaderboard / n/a
Avg benchmark score 69.8
Perf / dollar 587.8
Terminal-Bench 2.1 84.3
Humanity's Last Exam 55.3
AutoBench 48.8
Independent composite scores each on its own scale — not part of the average above
AA Intelligence Index 57.5
AA Agentic Index 58.2
AA-Omniscience Index (−100–100) 7.5
GDPval-AA v2 1763.8 ELO
LMSYS Chatbot Arena 1469.4 ELO
Providers
Available from
Z.AI — $0.075/$0.250/Mtok
Alibaba — $0.150/$0.470/Mtok

Cost at scale

1M tokens · 50/50 input/output
Projected cost of GLM-5.3-Flash vs Qwen3.8-Flash at increasing token volumes
VolumeGLM-5.3-FlashQwen3.8-FlashSavings
1M tokens $0.12 $0.23 $0.11 (47.8%)
10M tokens $1.19 $2.3 $1.11 (48.3%)
100M tokens $11.88 $23 $11.13 (48.4%)
1000M tokens $118.75 $230 $111.25 (48.4%)

When to pick which

distilled from the pricing and spec data above
Pick GLM-5.3-Flash if…
  • cost dominates: $0.119/Mtok blended vs $0.230 — 48.4% less on the same 50/50 token mix
  • your workload re-reads context (agents, RAG, long chats): cached input costs $0.015/Mtok — 80% off its list input price, a discount Qwen3.8-Flash doesn't offer
  • you want open weights — self-host it, fine-tune it, or exit the API entirely; Qwen3.8-Flash is closed
Pick Qwen3.8-Flash if…
  • this pairing gives Qwen3.8-Flash no edge on price, context, cache discount, or measured performance — pick it only for qualitative fit (ecosystem, compliance, model behavior on your prompts)
Try GLM-5.3-Flash on Z.AI

The cheaper option here — GLM-5.3-Flash costs $0.119/Mtok blended on Z.AI.

Get API key →

Summary

GLM-5.3-Flash by Z.AI costs $0.075/Mtok input and $0.250/Mtok output, with a 1M-token context window. It supports text, image input.

Qwen3.8-Flash by Alibaba costs $0.150/Mtok input and $0.470/Mtok output, with a 1M-token context window. It supports text input.

On a blended cost basis, GLM-5.3-Flash is 48.4% cheaper than Qwen3.8-Flash.

The two aren't directly comparable on average benchmark score: GLM-5.3-Flash has published per-benchmark results, while Qwen3.8-Flash does not yet.

Note: Pricing is per million tokens. Actual costs vary with usage patterns, prompt caching, and batch discounts. Always verify against official provider pricing pages.

More comparisons

pairs sharing a model with this page