On September 24, Google officially announced the launch of Gemini 3.8 Live with the Live Avatar feature, bringing near-real-time visual presentation to its native real-time conversation model. The feature combines low-latency video generation with speech to create an enterprise-grade interactive experience that can listen, observe, and communicate through a dynamic visual persona; it is now available on Gemini Enterprise.
Conversation Is Inherently Multimodal: A "Face" for the Enterprise Agent
Google said in its official blog: "Conversation itself is multimodal: we listen, observe, speak, and use facial expressions to communicate." Live Avatar unifies voice, vision, and facial expressions within an enterprise agent, making digital interaction richer and more natural through precise lip sync, natural expressions, and smooth turn-taking in conversation.
Live Avatar processes visual and audio input simultaneously to generate a richer conversational experience. Its core application scenarios include customer service and interactive guided tours, aiming to upgrade AI agents from "voice only" to "conversational partners with a face and expressions."
Asynchronous Tool Calling: Talking and Working in Parallel
Beyond visual presentation, Live Avatar leverages Gemini's advanced reasoning capabilities to support asynchronous tool calling. This means the model can trigger tool calls and retrieve data in the background while keeping the conversation going, so users do not experience a "silent wait" when handling complex tasks.
Google gives an example: you can ask Live Avatar to check you in at a hotel while continuing to discuss other matters with it—it queries the booking system in the background without interrupting the conversation.
Seamless Switching Across 97 Languages: Lip Sync and Expressions Adapt in Sync
Live Avatar supports native multilingual speech-to-speech synchronization, dynamically adapting lip sync and expressions, and switching seamlessly among 97 languages without reducing video fidelity or causing visual drift. This means that when the language switches mid-conversation, lip movements and facial animation must keep up in real time to maintain visual consistency.
Presets and Customization: A Two-Tier Approach to Brand Identity
Enterprises can choose a suitable persona from a diverse library of preset avatars, or generate a custom virtual persona from a single high-quality reference image, while preserving the reference subject's likeness, brand style, or character traits.
However, custom avatar generation is currently available only to enterprise allowlisted users, and requires contacting a Google sales representative to apply and allocate resources.
SynthID Watermark: Transparency and Security in Equal Measure
All audio and video output generated by Live Avatar embeds an invisible SynthID watermark, woven directly into the audio and video tracks, to help identify AI-generated content and reduce misinformation and misattribution. This safety design reflects Google's emphasis on respecting identity and content transparency.
Competitive Landscape: The AI Virtual Human Race Heats Up
Google's entry into the AI virtual human market puts it in direct competition with specialized platforms such as HeyGen and Synthesia. HeyGen's LiveAvatar platform likewise offers real-time AI avatar technology, supporting FULL mode and LITE mode, and can train a custom avatar from as little as a 2-minute video, priced at about $0.18/minute. Synthesia is known for enterprise training videos and is valued at $4 billion.
Unlike these "asynchronous" tools, Live Avatar is built for real-time interaction, a harder engineering problem and a different market.
The significance of Live Avatar is not that "AI has a face," but the architectural innovation of asynchronous tool calling—it lets AI "work" while conversing, upgrading interaction from "request-response" to "continuous companionship." When AI can chat with you while checking you in and querying data in the background, the boundaries of an enterprise agent's value are significantly broadened. Real-time lip sync across 97 languages provides multinational enterprises with a unified visual service interface, greatly reducing the localization costs of multilingual deployment.