The 2026 World Artificial Intelligence Conference (WAIC) and the High-Level Meeting on Global AI Governance are currently taking place in Shanghai from July 17th to 20th.
At the Tencent AI Application Innovation Forum held today, Huya Inc. CEO Huang Junhong unveiled "Huya VAM 1.0," a real-time multimodal digital human foundational model based on the DiT architecture.
During the demonstration, the model automatically generated several digital anchorpersons. Their voices and micro-expressions were so highly similar to real humans that the audience found them "extremely realistic."
It is reported that this model can generate an AI digital human capable of speaking, singing, and dancing from just a single photo, and it supports real-time interaction. The model outputs at a resolution of 480×832 and 28 frames per second in a real-time streaming format, capable of running continuously for over 24 hours. It possesses full-state human-like interactive simulation capabilities covering silent, listening, and speaking states, and supports full-duplex interaction, instant interruption, and natural transitions.
Comments