LIVE Cheapest paid: Granite 4.0 Micro $0.017/Mtok in 174 models tracked Updated Jul 25, 2026
Jul 25, 2026
ModelPriceWatch$/Mtok
Pricing / Compare / GLM-5.2 vs Qwen3.7-Max

GLM-5.2 vs Qwen3.7-Max

Side-by-side comparison of API pricing, specs, benchmarks, and capabilities

GLM-5.2 is 42% cheaper on blended cost ($2.90 vs $5.00/Mtok)
 
Zby Z.AI
2 providers: Together $1.40/$4.40 Z.AI $1.40/$4.40
Qwen3.7-Max Intro price
2 providers: Alibaba $2.50/$7.50 Together $1.25/$3.75
Overview
StatusCurrent flagship Open weights Current flagship
Released Jun 13, 2026 May 20, 2026
Pricing per million tokens
Input $1.40/Mtok $2.50/Mtok
Output $4.40/Mtok $7.50/Mtok
Blended avg $2.90/Mtok $5.00/Mtok
Cached input $0.260/Mtok $0.250/Mtok
Specifications
Context window 1M tokens 1M tokens
Parameters Proprietary Proprietary
Speed (TPS)
Modalities
Input
text
text
Benchmarks sources: Vellum LLM Leaderboard / Artificial Analysis, Alibaba, Qwen model card
GPQA Diamond 91.2 68
SWE-Bench Verified 45 45
Humanity's Last Exam 54.7 25
ARC-AGI 2 30 28
AIME 2025 70 70
MMMLU 80 80
BFCL 70 68
HumanEval 84 84
MATH 500 76 76
Avg benchmark score 76
Perf / dollar 26.2
Terminal-Bench 2.1 81
MCP Atlas 77
AIME 2026 99.2
Independent composite scores each on its own scale — not part of the average above
AA Intelligence Index 51.1 46
AA Agentic Index 43.1 30.6
AA-Omniscience Index (−100–100) 4 14.1
GDPval-AA v2 1510.3 ELO 1271.3 ELO
LMSYS Chatbot Arena 1469.2 ELO
Providers
Available from
Together — $1.40/$4.40/Mtok
Z.AI — $1.40/$4.40/Mtok
Alibaba — $2.50/$7.50/Mtok
Together — $1.25/$3.75/Mtok

Cost at scale — 1M tokens (50/50 input/output)

VolumeGLM-5.2Qwen3.7-MaxSavings
1M tokens $2.9 $5 $2.1 (42%)
10M tokens $29 $50 $21 (42%)
100M tokens $290 $500 $210 (42%)
1000M tokens $2900 $5000 $2100 (42%)

Summary

GLM-5.2 by Z.AI costs $1.40/Mtok input and $4.40/Mtok output, with a 1M-token context window. It supports text input and is available from 2 providers.

Qwen3.7-Max by Alibaba costs $2.50/Mtok input and $7.50/Mtok output, with a 1M-token context window. It supports text input and is available from 2 providers.

On a blended cost basis, GLM-5.2 is 42% cheaper than Qwen3.7-Max.

The two aren't directly comparable on average benchmark score: GLM-5.2 has published per-benchmark results, while Qwen3.7-Max does not yet — it is measured today on independent composites (see the table above).

Note: Pricing is per million tokens. Actual costs vary with usage patterns, prompt caching, and batch discounts. Always verify against official provider pricing pages.