Hy4 preview
by Tencent · 770B (A49B) parameters
Today's price · per 1M tokens
Input
$0.834
per 1M tokens
Output
$2.50
per 1M tokens
Blended
$1.25
blended $/1M — 3:1 weighted input:output
Cached input
$0.042
5% of input — prompt caching
How this price is scoped: Tencent lists this model in RMB on its TokenHub rate card — ¥6 in / ¥18 out / ¥0.3 cache-hit per 1M tokens. Those two region tabs are different rate cards in general (新加坡 reprices GLM-5.3 from ¥8/¥28/¥2 to ¥10.07524/¥31.66504/¥1.871116), but the Hy4 preview row is byte-identical on both, so this model has no separate international list. The USD above is that list at 7.197 CNY/USD, the rate implied to three figures by all three prices the FP8 endpoint Tencent itself operates quotes in USD ($0.834 / $2.501 / $0.042, the only endpoint serving this model). That is Tencent's own commercial rate, not a market snapshot: ECB spot on 2026-08-27 was 6.7203 CNY/USD, which would instead give $0.893 / $2.679 / $0.045. We publish what Tencent charges — FX_CURRENCY_POLICY rule (1) — not a spot re-conversion of its domestic card.
Price receipt
We read Tencent's own pricing page on Aug 28, 2026 and found Hy4 preview listed with a price on the same row — that read is where the number above comes from, recorded as a new model. The page prices it in CNY; the USD above is converted from that list at a dated FX rate — see the price note for the figures. We keep the page text we read; its content hash is 38746ea734.
Overview
Tencent Hy4 preview — Tencent's new-generation flagship Mixture-of-Experts model, 770B total parameters with 49B activated per token across 78 layers (256 routed experts plus one shared, top-8 routing), using Gated DeepSeek Sparse Attention with IndexCache and identity Hyper-Connections. Tencent open-sourced the instruct and FP8 weights under Apache 2.0 on 2026-08-27 and sells the model per token on its Tencent Cloud TokenHub rate card, serving it itself as an FP8 endpoint. Tencent's own model list gives a 1M context window with a 960K maximum input and a 64K maximum output; the endpoint Tencent operates reports 1,048,576 tokens, which is the figure published here. "preview" is Tencent's own name for this release, not a staging label of ours.
Deploy this open model on rented GPUs
Open weights mean you can self-host instead of paying per-token API prices. These platforms let you serve it on demand — often cheaper at scale.
Partner links — we may earn a commission if you sign up, at no cost to you. We only list platforms we'd recommend regardless.
Capabilities
struck through = not supportedBenchmark performance
accuracy % · higher is betterWe hold no per-benchmark accuracy scores for this model yet, so it has no accuracy average. That is a gap in our coverage, not a sign the model is untested — independent evaluators often publish a composite index for a new model long before releasing its per-benchmark numbers. The percentile above is its standing across the independent composites below.
- LMSYS Chatbot Arena: 1633.3 ELO (Human preference)
- Ranks #1 of 170 comparably-measured models by percentile score, across 1 independent measurement
- Ranks #33 of 170 comparably-tested models by normalized performance per dollar
Source: LMSYS Chatbot Arena (UC Berkeley) · updated Aug 28, 2026 · See full rankings →
Specifications
- Provider
- Tencent
- Context window
- 1M tokens
- Modality
- text
- Parameters
- 770B (A49B)
- Open source
- Yes — open weights available
- Released
- Aug 27, 2026
- Status
- Current
- Last updated
- Aug 28, 2026
- Tags
Availability verified: Aug 28, 2026 — listed on Tencent's own page