Unisound (09678) Upgrades Voice Foundation Models U2-ASR and U2-TTS with Expanded Multilingual Capabilities

Stock News07-28

Unisound (09678) has announced a comprehensive upgrade to its U2-ASR and U2-TTS voice models, strengthening the company's multimodal large model capabilities. This upgrade marks a significant expansion from covering hundreds of Chinese dialects to supporting a wide range of global languages, reinforcing the infrastructure for global voice interaction. The enhanced models are designed to provide efficient and low-barrier technical support for cross-border applications, embodied intelligence, and agent-based interactions.

In this upgrade, U2-ASR has added recognition capabilities for 13 new international languages, covering key overseas markets in Europe, Southeast Asia, the Middle East, and Latin America. Meanwhile, U2-TTS now supports speech synthesis for 8 additional Southeast Asian languages. As a result, the U2 voice large model now supports over 100 Chinese dialects and more than 15 international languages. This allows enterprises to process audio content in different languages by connecting to just one model and one set of APIs, significantly lowering the development, deployment, and maintenance threshold for multilingual voice services.

Under a unified evaluation benchmark, U2-ASR has demonstrated outstanding performance compared to industry-leading models, achieving an average character error rate (CER) of just 6.58% across 113 languages. In real-world business scenarios where language tags are not provided, the model maintains high accuracy through automatic language identification and closed-set routing, effectively avoiding language misidentification. Additionally, in the ChinaVoices Challenge 2026 Chinese multi-dialect speech recognition competition, U2-ASR achieved a top ranking with a CER of 9.235%, securing first place in the limited-data track and second place in the open-data track, further validating its advanced speech recognition technology.

U2-TTS has also demonstrated leading performance in subjective evaluations of intelligibility and naturalness compared to mainstream industry models. It utilizes a streaming neural network acoustic model to output 24kHz high-fidelity audio chunk by chunk in real time, significantly reducing initial packet latency and meeting the demands of strong real-time interactions such as live voice conversations. This upgrade enables seamless collaboration between U2-ASR and U2-TTS, providing enterprises with a complete multilingual voice interaction pipeline for building agent-based systems, positioning voice technology as a core interaction infrastructure for global business.

The upgraded U2-ASR and U2-TTS models are now fully available on the company's TokenHub large model service platform, with standard APIs open for access. The company will continue to adhere to its philosophy of "intelligence for good," leveraging technology to break down language barriers and promote the universal sharing of AI technology, ensuring that regions at different stages of development can equally enjoy the technological dividends of the AGI era.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Comments

We need your insight to fill this gap
Leave a comment