ARTICLE•NEWS 2+

Tencent Releases Hy4 Preview, A Major Leap for Hunyuan in Productivity-Class Models

28 August 2026•
Image for Tencent Releases Hy4 Preview, A Major Leap for Hunyuan in Productivity-Class Models

On August 28, 2026 Tencent Hunyuan officially released and open-sourced the latest large language model generation, Hy4 preview. As successor to Hy3, Hy4 preview brings a significant leap in scale, context length, and performance across real-world productivity tasks, from coding and office work to scientific research.

The release also confirms Hunyuan’s aggressive iteration rhythm since rebuilding infrastructure in February 2026. With a preview-first then official release approach, Tencent deliberately ships early to collect real-world feedback. Based on the official Tencent announcement on August 28 and the GitHub repository Tencent-Hunyuan/Hy4-preview, the fast feedback loop explains substantial maturation of Hy3 after the preview, and Hy4 preview represents the largest generation-over-generation gain recorded internally so far.

Key Specifications

Hy4 preview is built on a large-scale Mixture-of-Experts (MoE) architecture:

  • 770 billion total parameters with 49 billion active parameters per token
  • Context window supporting more than 1 million tokens
  • 78-layer backbone, first layer as standard dense FFN and remaining 77 layers as MoE with 256 routed experts and 1 shared expert per layer and top-8 active experts per token
  • 1 native MTP (Multi-Token Prediction) layer with 10 billion total and 0.7 billion active parameters to accelerate speculative decoding
  • Gated DeepSeek Sparse Attention (Gated DSA) with IndexCache plus identity Hyper-Connections (iHC) on residual paths
  • Hidden size 6144, 64 attention heads, vocab size 120832, and released under Apache License 2.0

Compared to Hy3 at 295 billion total and 21 billion active parameters, Hy4 preview is more than twice as large, with active parameters also doubling and context expanding from 256K to 1M. The scaling here does not serve headline chasing. The combination of Gated DSA and iHC reflects serious attention to architectural efficiency rather than mere parameter inflation. On Hugging Face and ModelScope, weights are available in both full and lighter FP8 variants for multi-GPU serving.

Focus on Real-World Productivity

Tencent positions Hy4 preview as a model designed specifically for productivity scenarios, including software engineering, office analytics, game development, and scientific research. Development involved collaboration with internal Tencent expert teams across domains such as software engineering, gaming, finance, and security, alongside deep co-design with products like CodeBuddy and WorkBuddy.

In an internal blind evaluation in the WorkBuddy environment with 163 experts and 203 software engineering tasks, Hy4 preview achieved an average score of 2.99 out of 4.00, slightly ahead of GLM 5.3 at 2.92 and Kimi K3 at 2.94. The gap remains narrow, so reading the result as on par rather than a decisive lead appears more accurate. Breaking the numbers down:

  • Against GLM 5.3, Hy4 won 46.8 percent, tied 12.8 percent, and lost 40.4 percent
  • Against Kimi K3, Hy4 won 51.2 percent, tied 7.9 percent, and lost 40.9 percent

Highlighted capabilities include:

Software Engineering

Stronger understanding, planning, debugging, and validation for long-context development tasks, including building complex frontend projects from scratch with improved visual taste.

Game Development

Generating playable prototypes from a single prompt up to full demos in Unity and continuing to refine complex projects over multiple turns.

Office and Analytics

Handling complex financial audits by filtering and analyzing many documents at once, then turning messy context into polished documents, spreadsheets, and presentations.

Scientific Research

Accelerating molecular dynamics simulations and beginning to build recursive self-improvement loops, spanning AI research, condensed matter physics, and pure mathematics.

Interestingly, Tencent openly acknowledges an early preview status with known limitations, namely a tendency to spend longer than needed when reasoning on complex tasks and a tendency to over-verify the model’s own results. Such candor rarely appears in flagship releases, and the openness adds credibility, while simultaneously signaling that the model is not yet fully ready for critical production workloads.

Hands-On Impressions From Direct Testing

Direct testing of Hy4 preview ran for several days via Tencent Cloud TokenHub and OpenRouter APIs using model ID tencent/hy4-preview, including trials of the free tier on WorkBuddy and CodeBuddy that the official announcement describes as open free for two weeks after launch, while Hy3 free access extends to September 30. Migration from Hy3 feels smooth, since recommended parameters remain temperature=0.9 and top_p=1.0, with reasoning defaulting to high for heavy tasks. For direct answers without long reasoning, adding chat_template_kwargs with reasoning_effort set to no_think suffices.

For daily coding, my initial impression places Hy4 as a noticeably more careful model than Hy3. When a refactor across a dozen files in my personal repository was requested, Hy4 did not jump straight to patches. Hy4 drafted a plan, checked dependencies, then executed and ran verification. Frontend builds generated from scratch also appeared more orderly, especially in spacing consistency and interaction polish, which aligns with one of Tencent’s stated focus areas.

In terms of scale, the gap from 295 billion on Hy3 to 770 billion on Hy4 shows up immediately in use, at least in my view. Capacity on layered tasks grows stronger, yet cost moves upward. API rates for Hy4 sit above Hy3, and multi-GPU demand for self-hosting remains high even on the FP8 variant. Although the pricing still looks affordable overall, my personal view is that the jump from Hy3 to Hy4 is noticeable, since inference cost per million tokens is genuinely higher for a far larger model. Balance between added capability and added expense becomes the central consideration when choosing Hy3 for fast economical work versus Hy4 for deep exploration.

On office work with many documents, Hy4 excels at turning messy context into structured artifacts. My trials with a bundle of financial reports and meeting notes in mixed formats showed Hy4 pulling key figures, building concise tables, and rewriting everything into a draft presentation without losing cross-file context. Yet the 1 million token context calls for careful use. Larger context does not guarantee perfect recall. More stable results emerged when core documents were filtered first instead of feeding everything in raw.

The over-verification pattern disclosed by Tencent also appeared during my testing. On complex tasks Hy4 sometimes spends extra steps re-checking an already correct result, which increases latency and token usage. A practical remedy involves defining a clear acceptance test in the prompt and instructing when to stop. For small fixes, the no-think mode is far more economical and sufficient. For heavy multi-file refactors, the high mode is where depth is truly needed.

Compared to my direct experience with Hy3, which felt underrated but remarkably careful on bug fixing, Hy4 reads as a more ambitious continuation. Hy3 delivered reliability through rare hallucination, Hy4 adds stronger drive on long-horizon agentic tasks. For careful daily work Hy3 remains solid, for layered exploration, long debugging, and multi-turn iteration Hy4 offers more headroom, provided cost and reasoning time receive careful management.

Benchmark Results

Hy4 preview was tested across multiple categories and compared with other top-tier models such as DeepSeek V4 Pro 0813, Qwen 3.8 Max, GLM 5.3, Kimi K3, GPT 5.6 Sol, and Claude Opus 5. Here are highlights from the Agentic Coding category, where gains are most pronounced:

BenchmarkHy3Hy4 preview
SWE-bench Multilingual75.882.9
SWE-bench Pro57.965.7
DeepSWE28.064.3
Terminal-Bench 2.170.885.4
CyberGym51.878.4
ProgramBench3.017.5
Harbor-Index15.639.6

On Terminal-Bench 2.1, Hy4 preview at 85.4 even surpasses some measurements for DeepSeek V4 Pro 0813 at 87.9 and 80.3* and sits close to other top-tier models. The most dramatic jump is on DeepSWE, rising from 28.0 on Hy3 to 64.3, more than doubling in a single generation. ProgramBench moving from 3.0 to 17.5 is also more than a fivefold improvement.

In other categories such as Agentic Search, Working Agent, STEM Agent, and Reasoning, Hy4 preview also shows consistent improvement over Hy3:

  • WideSearch up from 81.9 to 83.9
  • OneMillionBench with tools up from 51.5 to 65.4
  • Toolathlon Verified up to 74.1 and surpassing Qwen 3.8 Max and GPT 5.6 Sol
  • APEX Agents pass at 1 up to 37.1 and nearly matching Kimi K3 at 37.2
  • MathArena Apex 2025 up from 38.7 to 74.2, almost doubling
  • GPQA Diamond up from 90.9 to 92.3
  • SUPERChem up from 52.6 to 66.4

According to summaries of Hunyuan’s official blog and financial coverage on August 28, Hy4 preview was never the lowest scorer on any of 12 benchmarks, while total and active parameters remain far smaller than some competitors in the top tier. That pattern points to efficiency, not just scale.

Overall, Hy4 preview manages to close the gap with top-tier models like GPT 5.6 Sol and Claude Opus 5 on many benchmarks, although trailing remains on some metrics. Worth noting, most comparison figures marked with an asterisk (*) are results from Tencent’s own internal testing of competitor models, not official numbers published by each vendor. That does not mean the results are invalid, but waiting for independent evaluations before treating claims like surpassing DeepSeek V4 Pro as final remains advisable.

Pricing and Availability

On pricing, Hy4 preview is quite affordable:

  • USD 0.834 per million input tokens or about 6 yuan per million tokens
  • USD 2.501 per million output tokens or about 18 yuan per million tokens
  • USD 0.042 per million tokens for cache hit or about 0.3 yuan per million tokens

The pricing sits far below top-tier closed-source models like GPT 5.6 Sol or Claude Opus 5, and aligns with the common pattern of China-origin models such as DeepSeek, Qwen, GLM, and Kimi competing on cost efficiency. For developers with scale-sensitive inference costs, the pricing likely represents one of the main attractions, not just benchmark scores.

The model is already available via Tencent products such as WorkBuddy and CodeBuddy in both domestic and international versions, Yuanbao, and ima, and via API through Tencent Cloud TokenHub and OpenRouter. Model weights are also openly available on Hugging Face, ModelScope, GitCode, and CNB, with official deployment guides for vLLM and SGLang that expose an OpenAI-compatible API.

For self-hosting, Tencent provides ready-to-use images vllm/vllm-openai:hy4-preview and lmsysorg/sglang:hy4-preview with speculative decoding via MTP. Still, a 770 billion parameter backbone remains demanding even in FP8 and requires substantial multi-GPU capacity. For most teams, starting from a hosted API makes more sense than jumping straight to self-hosting.

Conclusion

Hy4 preview delivers a clear leap over Hy3, with 770 billion parameters, a 1 million token context, and a more careful working style on long tasks. In my personal view, the model is most convincing when used for layered exploration and complex code fixes, while Hy3 remains the economical choice for daily work. Although per-token pricing rises, the value stays affordable for teams that need the extra capacity.