Tencent Quietly Drops a Game-Changing AI Model, Matching Top Global Rivals Across Multiple Benchmarks—Free for Two Weeks

Deep News08-28 20:52

On August 28, Tencent's Hunyuan division unveiled and open-sourced its next-generation large language model, Hy4 preview. The new model boasts a staggering 770 billion total parameters, with 49 billion activated parameters, and a context window extending to 1 million tokens. Compared to the previous generation's 295 billion total parameters, this represents a massive leap in scale.

But sheer parameter count is only part of the story—Tencent also shared updated test results that show notable improvements across the board. In the Terminal Bench 2.1 evaluation, Hy4 preview scored an impressive 85.4 points, a full 14.6 points higher than its predecessor. According to Tencent's data, this result surpasses DeepSeek V4 Pro and ties with Claude Opus 5. The DeepSWE benchmark, which measures software engineering capabilities, saw an even more dramatic jump: Hy3 scored just 28.0 points, while Hy4 preview surged to 64.3 points. This metric goes beyond simple code generation, testing the model's ability to handle complete software development workflows.

Tencent also announced the API pricing for Hy4 preview: 6 yuan per million tokens for input, 18 yuan per million tokens for output, and a remarkably low 0.3 yuan per million tokens for cache hits. The model is now open-sourced and already integrated into several Tencent products, including Tencent Yuanbao, ima, WorkBuddy, and CodeBuddy. Developers can also access the API through Tencent Cloud's TokenHub and OpenRouter platforms. Notably, WorkBuddy is offering a limited-time two-week trial for Hy4 preview.

Zhang Jun, Tencent's Director of Public Relations, confirmed on August 28 that Hy4 preview is available for a limited-time trial in WorkBuddy starting immediately. Meanwhile, the free trial period for Hy3 has been extended until September 30. For everyday users, this setup is refreshingly straightforward—the new model is ready to test, while the older one remains accessible for continued use.

Tencent is highlighting four key application scenarios for Hy4 preview: software engineering, office analytics, game development, and scientific research. In the software engineering arena, Hy4 preview has enhanced its ability to understand, plan, debug, and validate long-cycle development tasks. Front-end code generation has also been optimized, with improvements to visual rendering and user interaction experience. For office applications, Tencent emphasizes complex document handling, data analysis, and cross-file collaboration—users can feed the model a pile of materials, and it will extract key information before generating documents, spreadsheets, and presentations.

In game development, the model can generate playable game prototypes from simple requirements and supports multi-round iterations for project refinement, all while integrating with mainstream game engines. For scientific research, Tencent revealed a particularly noteworthy result: working alongside the Hyra research agent, Hy4 preview contributed to research on the three-dimensional Blaschke–Lebesgue geometry problem. Tencent reports the relevant volume lower bound was pushed from 0.380799 to 0.41104.

Tencent also conducted an internal blind test involving 163 expert participants across 203 engineering tasks. Hy4 preview earned an average score of 2.99 out of 4. In the same evaluation, GLM-5.3 scored 2.92 and Kimi K3 scored 2.94. In head-to-head comparisons, Hy4 preview won against GLM-5.3 with a 46.8% win rate and beat Kimi K3 with a 51.2% win rate. It's worth noting that these figures come from Tencent's internal testing, and the specific tasks and scoring methodology haven't been fully disclosed—so they should be treated as reference points rather than definitive verdicts.

Tencent acknowledges that Hy4 preview is still in its early stages, with room for optimization in the training process. In some complex tasks, the model can exhibit excessive self-validation behavior. Additionally, Hy4 preview is a pure language model without native multimodal generation capabilities—for image-related tasks, the platform automatically switches to other models and calculates credits according to their respective rules.

The model is now live across multiple Tencent products, with WorkBuddy offering the two-week limited-time trial and Hy3 free until September 30. While 770 billion parameters sounds impressive on paper, the real proof lies in hands-on experience. For coding, document processing, and data analysis tasks, practical performance will likely speak louder than raw numbers. Can this Hy4 preview upgrade truly close the gap between China's open-source large models and the world's top-tier AI systems? The answer may well depend on how everyday users put it to the test.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Comments

We need your insight to fill this gap
Leave a comment