Google Gemini 3.8 Live Launches "Digital Human" Real-Time Conversations with Multilingual Support and Lip Sync

Deep News09-25 02:11

Google Cloud's Gemini 3.8 Live with Live Avatar feature has officially opened to enterprise users, integrating real-time video avatars with multilingual voice conversation capabilities, marking a shift in enterprise conversational AI from pure voice interaction toward multimodal visual interaction.

The feature officially launched this week within Gemini Enterprise, following a preview demonstration at Google Cloud Next 2026.

The new version supports real-time lip synchronization for video avatars, automatic recognition and switching across 97 languages, as well as backend tool invocation and API execution, while offering dual-region endpoint deployment across the United States and the European Union to meet enterprise compliance and data governance requirements.

In terms of customer adoption, Cox Automotive has already applied the technology to its Autotrader platform, building a car-buying AI assistant that can highlight screen content in real time and guide consumers through vehicle search and financing processes; Indian AI company Equal AI stated that, based on Gemini 3.8 Live, its platform now handles over one million real-time calls per day, covering nine Indian languages.

This marks Google's second AI model release this week. On Wednesday, Google announced the launch of two text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS.

As of Thursday's press time, Alphabet (GOOGL) was up 0.66% in intraday trading.

Core Capabilities: Multimodal Fusion Beyond Basic Voice

Gemini 3.8 Live with Live Avatar introduces a visual interaction layer on top of its existing voice foundation, primarily targeting multiple deployment scenarios including web, mobile, and physical interactive kiosks.

In terms of conversational continuity, the model adopts a native speech-to-speech architecture, supporting natural interruption and recovery without losing conversation context or interrupting background transaction processing.

Regarding tool invocation capabilities, the model can asynchronously execute tool calls and API requests in the background while maintaining continuous conversation, making the interaction experience smoother without noticeable pauses caused by background tasks.

In terms of visual understanding, Live Avatar supports synchronized processing of real-time camera feeds and screen sharing, which can be analyzed in parallel with audio input, providing richer contextual awareness for service scenarios.

On language coverage, support for 97 languages with automatic detection gives it the potential for scaled deployment across global multilingual markets.

Trust Mechanisms: Watermarking Technology and Identity Controls in Parallel

In response to deepfake and identity misuse risks, Google has adopted a dual-layer control strategy in the compliance design of this feature.

At the avatar usage level, enterprise users can deploy from Google's library of pre-built avatars; the custom avatar feature, however, has strict enterprise whitelisting and identity verification mechanisms, and is currently limited to enterprises whose applications have been approved. At the content traceability level, all generated audio and video streams are embedded with Google's SynthID imperceptible watermark to ensure that AI-generated content remains identifiable and verifiable throughout its distribution.

The introduction of this mechanism, to a certain extent, responds to the regulatory compliance pressures enterprises face when deploying conversational AI, especially in heavily regulated industries such as financial services, healthcare, and customer service.

Customer Deployments: From Car-Buying Assistants to Million-Scale Call Scenarios

Several companies have already disclosed progress on applications based on Gemini 3.8 Live, spanning multiple verticals including automotive e-commerce, telecommunications customer service, and CRM.

Cox Automotive Chief Product Officer Marianne Johnson stated that Autotrader's conversational AI avatar combines natural language descriptions with inventory matching, representing an important step in realizing its "connected intelligence" vision, which aims to enable every consumer interaction to draw on Cox Automotive's full data assets.

Equal AI Chief Executive Officer Akhilesh Damaraju said that, leveraging Gemini 3.8 Live's improvements in interruption handling, multilingual conversation, and tool invocation stability, the platform now processes over one million real-time calls per day, covering nine Indian languages, with the goal of building a personal AI that "knows you, speaks your language, and is always on your side."

Salesforce Agentforce Product Vice President Bob Van Osten stated that the collaboration between Salesforce AI Research and Google aims to explore the combination of real-time multimodal capabilities and agentic AI to create a customer service experience spanning the entire journey from first contact to problem resolution.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Comments

We need your insight to fill this gap
Leave a comment