Tencent Hy4 Preview: 770B MoE Architecture and 1M Context Window
Detailed analysis of Tencent Hy4 preview: 770B MoE architecture (49B active), 1M context window, Gated DSA attention, benchmark scores, and Apache 2.0 weights.


Synthesizing article benchmarks & model metrics...
Tencent has officially launched Hy4 preview, a massive 770-billion-parameter Mixture-of-Experts (MoE) foundation model designed for long-horizon autonomous coding, enterprise productivity, agentic workflows, and scientific computing.
Released under the Apache License 2.0, Hy4 preview represents one of the largest open-weight models to date, combining a 1,000,000-token context window with a sparse architecture that activates only 49B parameters per token (6.36% activation ratio). The release includes both full-precision BF16 base weights and optimized FP8 checkpoints across public repositories including Hugging Face and ModelScope. Compare where Hy4 ranks against global competitors in our AI Model Leaderboard and Full LLM Model Directory.
+---------------------------------------------------------------+| Tencent Hy4 Preview Overview |+-------------------+-------------------------------------------+| Total Parameters | 770 Billion Backbone Parameters || Active Parameters | 49 Billion per Token (8 Routed + 1 Shared)|| Context Capacity | 1,048,576 Tokens (1M Context Window) || Layer Count | 78 Layers (1 Dense FFN + 77 MoE Blocks) || Attention Design | Gated DeepSeek Sparse Attention (DSA) || License | Permissive Apache License 2.0 |+-------------------+-------------------------------------------+Architectural Breakdown: 770B Capacity with 49B Sparsity
The engineering design of Hy4 preview decouples total model capacity from per-token compute demands through fine-grained expert routing:
| Architectural Component | Specification | Operational Design |
|---|---|---|
| Total Layers | 78 | 1 Initial Dense FFN Layer + 77 MoE Transformer Blocks |
| Expert Allocation | 256 Routed | 8 Routed Experts activated per token + 1 Shared Expert |
| Attention Scheme | Gated DSA | Gated DeepSeek Sparse Attention with IndexCache |
| Speculative Decoding | MTP Layer | Native Multi-Token Prediction for accelerated generation |
| Residual Routing | Hyper-Connect | Identity Hyper-Connections across deep layers |
| Precision Formats | BF16 & FP8 | Native FP8 quantization support for distributed serving |
+------------------------------------------------------------------+| Hy4 MoE Layer Routing Topology || || Input Token Vector || | || v || +--------------+ +------------------------------------+ || | MoE Router | ---> | 1 Shared Expert (Always Active) | || +--------------+ +------------------------------------+ || | || +------------> [ Top-8 of 256 Routed Experts ] || | || v || Accumulated Token Output (49B Active) |+------------------------------------------------------------------+Key Technical Innovations
- Gated DeepSeek Sparse Attention (Gated DSA): Employs dynamic gating on top of sparse attention to prune irrelevant token connections in long documents while maintaining semantic continuity.
- IndexCache Layer Reuse: Caches and reuses attention index maps across transformer layers, drastically reducing KV-cache bandwidth requirements across 1M-token contexts.
- Multi-Token Prediction (MTP): Includes dedicated speculative prediction heads to emit multiple tokens per forward step, accelerating decoding throughput.
- Self-Optimized Infrastructure: Tencent reports utilizing early Hy4 checkpoints to profile kernel bottlenecks and automate operator fusion, achieving a 31.8% improvement in end-to-end serving throughput during internal testing.
Long-Context Capability: 1M-Token Context Window
The 1-million-token context window allows developers to feed entire multi-repository systems, large technical specifications, and longitudinal conversation logs directly into prompt memory.
Target Workload Domains
- Enterprise Codebase Refactoring: Ingesting entire code repositories with multiple dependency graphs for cross-file refactoring.
- Multi-Document Synthesis: Cross-referencing hundreds of financial filings, legal contracts, or scientific papers in a single inference call.
- Autonomous Agent Persistence: Retaining complete execution histories, terminal outputs, and error recovery traces across multi-hour agent runs.
- Scientific Research: Simulating complex multi-step pipelines across molecular dynamics, condensed matter physics, and pure mathematics.
Benchmark Performance & Evaluation
Tencent published evaluation data comparing Hy4 preview against leading frontier models, including Qwen 3.8 Max, DeepSeek V4 Pro 0813, GPT-5.6 Sol, GLM-5.3, Kimi K3, and Claude Opus 5.
Standardized Benchmark Scores
| Benchmark Suite | Hy4 Preview | Category Leader Score | Leader Margin Analysis | Primary Capability Tested |
|---|---|---|---|---|
| Terminal-Bench 2.1 | 80.3 | 88.3 (Kimi K3) | Kimi leads by +8.0 pts (+10.0%) | Bash & CLI command execution |
| Toolathlon-Verified | 74.1 | 76.5 (Leader) | Leader leads by +2.4 pts (+3.2%) | Multi-tool orchestration & chaining |
| OneMillionBench (Tools) | 65.4 | 68.1 (Leader) | Leader leads by +2.7 pts (+4.1%) | Long-context tool reasoning (1M tokens) |
| DeepSWE | 64.3% | 74.7% (DeepSeek V4 Pro) | DeepSeek leads by +10.4 pts (+16.2%) | Real-world GitHub issue resolution |
| SWE Atlas Refactoring | 53.3% | 60.0% (Leader) | Leader leads by +6.7 pts (+12.6%) | Codebase-wide refactoring |
| BioMysteryBench | 71.3% | 73.1% (Leader) | Leader leads by +1.8 pts (+2.5%) | Complex biological & biomedical reasoning |
| Humanity’s Last Exam | 43.4% | 53.2% (Leader) | Leader leads by +9.8 pts (+22.6%) | Frontier academic reasoning |
| PostTrainBench | 35.6% | 36.2% (GLM-5.3) | GLM-5.3 leads by +0.6 pts (+1.7%) | Post-training capability retention |
| APEX-Agents (Pass@1) | 37.7% | 41.8% (Leader) | Leader leads by +4.1 pts (+10.9%) | High-difficulty autonomous agent planning |
| Agent’s Last Exam (ALE) | 22.8% | 27.6% (Leader) | Leader leads by +4.8 pts (+21.1%) | Long-horizon multi-step planning |
| ProgramBench | 17.5% | 25.0% (Leader) | Leader leads by +7.5 pts (+42.9%) | Algorithmic program synthesis |
| HorizonMath (Pass@4) | 8.8% | 10.6% (Leader) | Leader leads by +1.8 pts (+20.5%) | Frontier mathematical theorem proving |
For deep dives into these evaluation methodologies, explore our Humanity’s Last Exam Guide and SWE-bench Verified Explainer.
Blind Expert Evaluation (Human Engineering Tasks)
Tencent conducted an internal blind evaluation involving 163 human software experts across 203 real-world engineering tasks (evaluated on a 1–4 scale):
| Comparison Metric | Hy4 Preview | GLM-5.3 | Kimi K3 |
|---|---|---|---|
| Average Score | 2.99 / 4.0 | 2.92 / 4.0 | 2.94 / 4.0 |
| Win Rate vs. Competitor | – | 46.8% | 51.2% |
| Tie Rate | – | 12.8% | 7.9% |
| Loss Rate | – | 40.4% | 40.9% |
Methodological Note: Internal blind evaluations and vendor-reported benchmarks should be validated through neutral, independent test harnesses.
Local Deployment and Hardware Considerations
Because Hy4 preview is licensed under Apache 2.0, organizations can host the model on private infrastructure without proprietary API dependencies.
Serving Engines & Integration
- Recommended Engines: vLLM and SGLang.
- Quantization Options: Native FP8 weights are provided alongside BF16 base weights.
- Tool Protocol: Native support for OpenAI-compatible function calling, JSON schema outputs, and structured reasoning parsers.
Hardware Infrastructure Requirements
Deploying a 770B-parameter MoE model requires substantial hardware resources despite its 49B active per-token footprint:
- FP8 Serving: Minimum 8x 80GB GPUs (e.g., NVIDIA H100/A100 clusters) utilizing Tensor Parallelism (TP=8).
- BF16 Full Precision: Typically requires 16x 80GB GPUs across multi-node setups to accommodate both model weights and large KV caches at 1M context.
Known Limitations and Operational Nuances
Tencent identifies several operational characteristics to consider during production pilots:
- Reasoning Verbosity: The preview build tends to generate longer internal reasoning traces than strictly necessary for simple tasks.
- Over-Verification Cycles: The model occasionally executes redundant verification passes before returning final answers, increasing end-to-end latency.
- Inference Latency Trade-offs: While active parameters are kept to 49B, memory routing across 256 experts requires high-bandwidth GPU interconnects (NVLink).
How to Test Hy4 Preview
Tencent provides multiple entry points for evaluating the model:
- Free Web UI Testing: Available free on Tencent WorkBuddy and CodeBuddy for two weeks post-launch.
- Hosted API Gateways: Accessible via Tencent Cloud TokenHub and OpenRouter.
- Open Weights Repositories: Available for download on Hugging Face, ModelScope, and GitCode.
Final Verdict
Tencent Hy4 preview establishes a new high-water mark for open-weight MoE architecture.
By combining 770B total capacity, 49B active compute (6.36% activation ratio), 1M context depth, and Apache 2.0 licensing, Tencent offers developers a transparent alternative to closed frontier models. While competing proprietary models maintain leads in certain coding and mathematical benchmarks, Hy4 preview’s strong showing on Terminal-Bench (80.3) and Toolathlon (74.1) makes it an essential candidate for engineering teams building self-hosted, long-context agent pipelines.
Explore comparative rankings in our AI Model Leaderboard and run head-to-head simulations in the Compare Arena.
Benchmark and pricing data is aggregated from OpenRouter, Artificial Analysis, and models.dev, then scored with the published RankLLMs Index. These are the primary sources behind the numbers in this article.
- •Tencent Official Announcement: Release of Open-Source Tencent Hy4 Preview(Primary Source →)
- •GitHub: Tencent-Hunyuan/Hy4-preview Repository & Deployment Guide(Primary Source →)
- •Hugging Face: Tencent Hunyuan Hy4-preview Model Weights(Primary Source →)

Lucky Yaduvanshi(luckyyaduvanshi.in →)
Founder of RankLLMs • AI Researcher & Software Engineer focusing on LLM benchmarking, DevOps, and autonomous coding agents.
Recommended Reading
Continue exploring related technical benchmarks, model deep-dives, and autonomous coding agent guides.

Alibaba Model Studio Token Plan: Pricing Structure and Claude Code Integration
Comprehensive review of the Alibaba Cloud Model Studio Token Plan. We break down Singapore region access, Credits math, 7-day rolling limits, night discounts, Reddit community benchmarks, and Claude Code setup.
Lucky Yaduvanshi
GLM-5.3 Flash vs Muse Spark 1.2 Contributor: Command Code Evaluation
Comparing GLM-5.3 Flash and Meta's Muse Spark 1.2 Contributor in Command Code: coding speed, unit test accuracy, and developer plan economics.
Lucky Yaduvanshi
Ox Alpha: How Z.ai Stealth-Tested GLM-5.3-Flash in Production
Z.ai confirms mystery model Ox Alpha was an unbranded preview of GLM-5.3-Flash. Discover 320B MoE specs, benchmark results, and what the MIT release means.
Lucky Yaduvanshi
Claude Sonnet 5: Permanent $2/$10 Pricing and Developer Impact
Anthropic has formalized Claude Sonnet 5's $2/$10 API pricing permanently, canceling scheduled increases. Here is the token economics breakdown.
Lucky Yaduvanshi