Today's price · per 1M tokens
Input
$2.00
per 1M tokens
Output
$6.00
per 1M tokens
Blended
$3.00
blended $/1M — 3:1 weighted input:output
How this price is scoped: Alibaba Model Studio International list price, single 0<Token≤1M tier, from the vendor's open-source Qwen table. The row carries no discount label, so this is the standard list rate. Alibaba marks the SKU as eligible for the context-caching discount but prints no cache rate in the table itself — its own gateway endpoint quotes $0.25 per 1M cached input tokens while the page's general footnote gives 10% of standard input as its example, so cached input is left unset here rather than picking one of the two.
Price receipt
We read Alibaba's own pricing page on and found qwen3.8-2.4t-a95b listed with a price on the same row — that read is where the number above comes from, recorded as a new model. We keep the page text we read; its content hash is ee2e068218.
Overview
The open-weights sibling of Alibaba's Qwen3.8 flagship — a 2.4T-parameter mixture-of-experts model with 95B active parameters, served by Alibaba itself on Model Studio at $2.00/$6.00 per 1M tokens, the same rate as the proprietary Qwen3.8-Max SKU. Text in, text out, with thinking and non-thinking modes billed alike. The weights were published on 2026-08-08 in Qwen's own Hugging Face org, and Alibaba's hosted 1M-context SKU appeared on its Model Studio rate card between 2026-08-12 and 2026-08-15.
Deploy this open model on rented GPUs
Open weights mean you can self-host instead of paying per-token API prices. These platforms let you serve it on demand — often cheaper at scale.
Partner links — we may earn a commission if you sign up, at no cost to you. We only list platforms we'd recommend regardless.
Capabilities
struck through = not supportedBenchmark performance
accuracy % · higher is betterEvery per-benchmark score we hold for this model is reported by its own vendor, so none is counted toward an accuracy average or any ranking. The percentile above is its standing across the independent composites below.
- AA Intelligence Index: 39.9 (Artificial Analysis composite across reasoning, knowledge and coding evals)
- AA-Omniscience Index: 4.3 (−100–100)
- GDPval-AA v2: 1627.8 ELO (Real-world work tasks, human baseline = 1000)
- Ranks #27 of 188 comparably-measured models by percentile score, across 3 independent measurements
- Ranks #115 of 188 comparably-tested models by normalized performance per dollar
23.9 tokens/sec output
Source: Artificial Analysis, Qwen3.8-2.4T-A95B Hugging Face model card (Qwen team) · updated · See full rankings →
Specifications
- Provider
- Alibaba
- Context window
- 1M tokens
- Modality
- text
- Parameters
- 2.4T (A95B)
- Open source
- Yes — open weights available
- Released
- Status
- Current
- Last updated
- Tags
Availability verified: — listed on Alibaba's own page