AI ModelsTencent HunyuanOpen Weight ModelsAI AgentsCoding AI

Tencent Hy4 Preview: 770B MoE Architecture and 1M Context Window

Detailed analysis of Tencent Hy4 preview: 770B MoE architecture (49B active), 1M context window, Gated DSA attention, benchmark scores, and Apache 2.0 weights.

Lucky Yaduvanshi
Lucky Yaduvanshi
Founder & AI Lead
Aug 29, 2026•Updated Sep 24, 2026•7 min read
Independent technical benchmark • Primary data & verified methodology cited below
Tencent Hy4 Preview 770B MoE Architecture and 1M Context Benchmark Analysis

Tencent has officially launched Hy4 preview, a massive 770-billion-parameter Mixture-of-Experts (MoE) foundation model designed for long-horizon autonomous coding, enterprise productivity, agentic workflows, and scientific computing.

Released under the Apache License 2.0, Hy4 preview represents one of the largest open-weight models to date, combining a 1,000,000-token context window with a sparse architecture that activates only 49B parameters per token (6.36% activation ratio). The release includes both full-precision BF16 base weights and optimized FP8 checkpoints across public repositories including Hugging Face and ModelScope. Compare where Hy4 ranks against global competitors in our AI Model Leaderboard and Full LLM Model Directory.

+---------------------------------------------------------------+
| Tencent Hy4 Preview Overview |
+-------------------+-------------------------------------------+
| Total Parameters | 770 Billion Backbone Parameters |
| Active Parameters | 49 Billion per Token (8 Routed + 1 Shared)|
| Context Capacity | 1,048,576 Tokens (1M Context Window) |
| Layer Count | 78 Layers (1 Dense FFN + 77 MoE Blocks) |
| Attention Design | Gated DeepSeek Sparse Attention (DSA) |
| License | Permissive Apache License 2.0 |
+-------------------+-------------------------------------------+

Architectural Breakdown: 770B Capacity with 49B Sparsity

The engineering design of Hy4 preview decouples total model capacity from per-token compute demands through fine-grained expert routing:

Architectural Component Specification Operational Design
Total Layers 78 1 Initial Dense FFN Layer + 77 MoE Transformer Blocks
Expert Allocation 256 Routed 8 Routed Experts activated per token + 1 Shared Expert
Attention Scheme Gated DSA Gated DeepSeek Sparse Attention with IndexCache
Speculative Decoding MTP Layer Native Multi-Token Prediction for accelerated generation
Residual Routing Hyper-Connect Identity Hyper-Connections across deep layers
Precision Formats BF16 & FP8 Native FP8 quantization support for distributed serving
+------------------------------------------------------------------+
| Hy4 MoE Layer Routing Topology |
| |
| Input Token Vector |
| | |
| v |
| +--------------+ +------------------------------------+ |
| | MoE Router | ---> | 1 Shared Expert (Always Active) | |
| +--------------+ +------------------------------------+ |
| | |
| +------------> [ Top-8 of 256 Routed Experts ] |
| | |
| v |
| Accumulated Token Output (49B Active) |
+------------------------------------------------------------------+

Key Technical Innovations

  1. Gated DeepSeek Sparse Attention (Gated DSA): Employs dynamic gating on top of sparse attention to prune irrelevant token connections in long documents while maintaining semantic continuity.
  2. IndexCache Layer Reuse: Caches and reuses attention index maps across transformer layers, drastically reducing KV-cache bandwidth requirements across 1M-token contexts.
  3. Multi-Token Prediction (MTP): Includes dedicated speculative prediction heads to emit multiple tokens per forward step, accelerating decoding throughput.
  4. Self-Optimized Infrastructure: Tencent reports utilizing early Hy4 checkpoints to profile kernel bottlenecks and automate operator fusion, achieving a 31.8% improvement in end-to-end serving throughput during internal testing.

Long-Context Capability: 1M-Token Context Window

The 1-million-token context window allows developers to feed entire multi-repository systems, large technical specifications, and longitudinal conversation logs directly into prompt memory.

Target Workload Domains

  • Enterprise Codebase Refactoring: Ingesting entire code repositories with multiple dependency graphs for cross-file refactoring.
  • Multi-Document Synthesis: Cross-referencing hundreds of financial filings, legal contracts, or scientific papers in a single inference call.
  • Autonomous Agent Persistence: Retaining complete execution histories, terminal outputs, and error recovery traces across multi-hour agent runs.
  • Scientific Research: Simulating complex multi-step pipelines across molecular dynamics, condensed matter physics, and pure mathematics.

Benchmark Performance & Evaluation

Tencent published evaluation data comparing Hy4 preview against leading frontier models, including Qwen 3.8 Max, DeepSeek V4 Pro 0813, GPT-5.6 Sol, GLM-5.3, Kimi K3, and Claude Opus 5.

Standardized Benchmark Scores

Benchmark Suite Hy4 Preview Category Leader Score Leader Margin Analysis Primary Capability Tested
Terminal-Bench 2.1 80.3 88.3 (Kimi K3) Kimi leads by +8.0 pts (+10.0%) Bash & CLI command execution
Toolathlon-Verified 74.1 76.5 (Leader) Leader leads by +2.4 pts (+3.2%) Multi-tool orchestration & chaining
OneMillionBench (Tools) 65.4 68.1 (Leader) Leader leads by +2.7 pts (+4.1%) Long-context tool reasoning (1M tokens)
DeepSWE 64.3% 74.7% (DeepSeek V4 Pro) DeepSeek leads by +10.4 pts (+16.2%) Real-world GitHub issue resolution
SWE Atlas Refactoring 53.3% 60.0% (Leader) Leader leads by +6.7 pts (+12.6%) Codebase-wide refactoring
BioMysteryBench 71.3% 73.1% (Leader) Leader leads by +1.8 pts (+2.5%) Complex biological & biomedical reasoning
Humanity’s Last Exam 43.4% 53.2% (Leader) Leader leads by +9.8 pts (+22.6%) Frontier academic reasoning
PostTrainBench 35.6% 36.2% (GLM-5.3) GLM-5.3 leads by +0.6 pts (+1.7%) Post-training capability retention
APEX-Agents (Pass@1) 37.7% 41.8% (Leader) Leader leads by +4.1 pts (+10.9%) High-difficulty autonomous agent planning
Agent’s Last Exam (ALE) 22.8% 27.6% (Leader) Leader leads by +4.8 pts (+21.1%) Long-horizon multi-step planning
ProgramBench 17.5% 25.0% (Leader) Leader leads by +7.5 pts (+42.9%) Algorithmic program synthesis
HorizonMath (Pass@4) 8.8% 10.6% (Leader) Leader leads by +1.8 pts (+20.5%) Frontier mathematical theorem proving

For deep dives into these evaluation methodologies, explore our Humanity’s Last Exam Guide and SWE-bench Verified Explainer.

Blind Expert Evaluation (Human Engineering Tasks)

Tencent conducted an internal blind evaluation involving 163 human software experts across 203 real-world engineering tasks (evaluated on a 1–4 scale):

Comparison Metric Hy4 Preview GLM-5.3 Kimi K3
Average Score 2.99 / 4.0 2.92 / 4.0 2.94 / 4.0
Win Rate vs. Competitor – 46.8% 51.2%
Tie Rate – 12.8% 7.9%
Loss Rate – 40.4% 40.9%

Methodological Note: Internal blind evaluations and vendor-reported benchmarks should be validated through neutral, independent test harnesses.

Local Deployment and Hardware Considerations

Because Hy4 preview is licensed under Apache 2.0, organizations can host the model on private infrastructure without proprietary API dependencies.

Serving Engines & Integration

  • Recommended Engines: vLLM and SGLang.
  • Quantization Options: Native FP8 weights are provided alongside BF16 base weights.
  • Tool Protocol: Native support for OpenAI-compatible function calling, JSON schema outputs, and structured reasoning parsers.

Hardware Infrastructure Requirements

Deploying a 770B-parameter MoE model requires substantial hardware resources despite its 49B active per-token footprint:

  • FP8 Serving: Minimum 8x 80GB GPUs (e.g., NVIDIA H100/A100 clusters) utilizing Tensor Parallelism (TP=8).
  • BF16 Full Precision: Typically requires 16x 80GB GPUs across multi-node setups to accommodate both model weights and large KV caches at 1M context.

Known Limitations and Operational Nuances

Tencent identifies several operational characteristics to consider during production pilots:

  1. Reasoning Verbosity: The preview build tends to generate longer internal reasoning traces than strictly necessary for simple tasks.
  2. Over-Verification Cycles: The model occasionally executes redundant verification passes before returning final answers, increasing end-to-end latency.
  3. Inference Latency Trade-offs: While active parameters are kept to 49B, memory routing across 256 experts requires high-bandwidth GPU interconnects (NVLink).

How to Test Hy4 Preview

Tencent provides multiple entry points for evaluating the model:

  • Free Web UI Testing: Available free on Tencent WorkBuddy and CodeBuddy for two weeks post-launch.
  • Hosted API Gateways: Accessible via Tencent Cloud TokenHub and OpenRouter.
  • Open Weights Repositories: Available for download on Hugging Face, ModelScope, and GitCode.

Final Verdict

Tencent Hy4 preview establishes a new high-water mark for open-weight MoE architecture.

By combining 770B total capacity, 49B active compute (6.36% activation ratio), 1M context depth, and Apache 2.0 licensing, Tencent offers developers a transparent alternative to closed frontier models. While competing proprietary models maintain leads in certain coding and mathematical benchmarks, Hy4 preview’s strong showing on Terminal-Bench (80.3) and Toolathlon (74.1) makes it an essential candidate for engineering teams building self-hosted, long-context agent pipelines.

Explore comparative rankings in our AI Model Leaderboard and run head-to-head simulations in the Compare Arena.

Sources, Disclosures & Primary Benchmark Data

Benchmark and pricing data is aggregated from OpenRouter, Artificial Analysis, and models.dev, then scored with the published RankLLMs Index. These are the primary sources behind the numbers in this article.

Share Article
Lucky Yaduvanshi

Lucky Yaduvanshi(luckyyaduvanshi.in →)

Founder of RankLLMs • AI Researcher & Software Engineer focusing on LLM benchmarking, DevOps, and autonomous coding agents.

Curated Research

Recommended Reading

Continue exploring related technical benchmarks, model deep-dives, and autonomous coding agent guides.

Explore All 38 Guides→