ModelPriceWatch.com
Last scan 2026-09-29 Models tracked 272 Providers 35 Cheapest paid Granite 4.0 H Micro $0.017/Mtok in Every price links to its source

Qwen3.8-2.4T-A95B

by Alibaba · 2.4T (A95B) parameters

Current open weights open weights mid tier

Today's price · per 1M tokens

Input

$2.00

per 1M tokens

Output

$6.00

per 1M tokens

Blended

$3.00

blended $/1M — 3:1 weighted input:output

How this price is scoped: Alibaba Model Studio International list price, single 0<Token≤1M tier, from the vendor's open-source Qwen table. The row carries no discount label, so this is the standard list rate. Alibaba marks the SKU as eligible for the context-caching discount but prints no cache rate in the table itself — its own gateway endpoint quotes $0.25 per 1M cached input tokens while the page's general footnote gives 10% of standard input as its example, so cached input is left unset here rather than picking one of the two.

Source: official Alibaba pricing · read MODELPRICEWATCH.COM · 2026-09-29

Price receipt

We read Alibaba's own pricing page on and found qwen3.8-2.4t-a95b listed with a price on the same row — that read is where the number above comes from, recorded as a new model. We keep the page text we read; its content hash is ee2e068218.

Overview

The open-weights sibling of Alibaba's Qwen3.8 flagship — a 2.4T-parameter mixture-of-experts model with 95B active parameters, served by Alibaba itself on Model Studio at $2.00/$6.00 per 1M tokens, the same rate as the proprietary Qwen3.8-Max SKU. Text in, text out, with thinking and non-thinking modes billed alike. The weights were published on 2026-08-08 in Qwen's own Hugging Face org, and Alibaba's hosted 1M-context SKU appeared on its Model Studio rate card between 2026-08-12 and 2026-08-15.

Run it yourself

Deploy this open model on rented GPUs

Open weights mean you can self-host instead of paying per-token API prices. These platforms let you serve it on demand — often cheaper at scale.

Partner links — we may earn a commission if you sign up, at no cost to you. We only list platforms we'd recommend regardless.

Capabilities

struck through = not supported
Input 1/4
Text ✓ Image Audio Video
Output
Text ✓
Features 2/9
Prompt caching Reasoning Coding Fast inference Long context ✓ Open weights ✓ Multimodal Web search Realtime

Benchmark performance

accuracy % · higher is better
Percentile vs all tracked models 74.8th3 independent measurements
Percentile per $/Mtok 24.9
GPQA Diamond vendor
92.6%
Terminal-Bench 2.1 vendor
86.6%
Humanity's Last Exam vendor
43.6%

Every per-benchmark score we hold for this model is reported by its own vendor, so none is counted toward an accuracy average or any ranking. The percentile above is its standing across the independent composites below.

Independent composite scores
  • AA Intelligence Index: 39.9 (Artificial Analysis composite across reasoning, knowledge and coding evals)
  • AA-Omniscience Index: 4.3 (−100–100)
  • GDPval-AA v2: 1627.8 ELO (Real-world work tasks, human baseline = 1000)
How it stacks up
  • Ranks #27 of 188 comparably-measured models by percentile score, across 3 independent measurements
  • Ranks #115 of 188 comparably-tested models by normalized performance per dollar

23.9 tokens/sec output

Source: Artificial Analysis, Qwen3.8-2.4T-A95B Hugging Face model card (Qwen team) · updated · See full rankings →

Specifications

Provider
Alibaba
Context window
1M tokens
Modality
text
Parameters
2.4T (A95B)
Open source
Yes — open weights available
Released
Status
Current
Last updated
Tags
open-weightsmoe

Availability verified: — listed on Alibaba's own page