Google made Gemini 3.8 Live with Live Avatar generally available in Gemini Enterprise on Thursday, giving its real-time conversational model an animated on-screen persona that lip-syncs and shifts expression as it talks. The feature couples the model's native speech-to-speech dialogue with low-latency streaming video generation, so the system listens, watches a camera feed and answers with a visible face rather than a waveform.
Key takeaways
- Live Avatar reached general availability in Gemini Enterprise on September 24, one week after the underlying Gemini 3.8 Live model shipped.
- The avatar maintains lip-sync and facial expression across 97 languages and can switch between them mid-conversation without what Google calls visual drift.
- Every generated audio and video stream carries an imperceptible SynthID watermark, and building a custom avatar from a reference photo requires enterprise allowlisting.
What Live Avatar actually generates
In its launch post from the Gemini audio team, Google frames the feature as native coupling rather than a bolt-on renderer: video generation and speech are produced together, which is what allows precise lip-sync and turn-taking instead of a mouth animation chasing an audio track.
The more practical addition is asynchronous tool execution. The avatar can fire a tool call and pull data in the background while the conversation keeps running, so a caller is not left staring at a frozen face during a lookup. Google's demo walks through a hotel check-in where the dialogue continues uninterrupted while function calls resolve behind the scenes.
Live visual understanding rounds it out. The model processes camera feeds and screen shares alongside audio simultaneously, which makes it a genuinely multimodal agent rather than a voice bot with a portrait attached. Google is positioning it for web, mobile and interactive kiosks.
How Google handled the identity problem
A model that can render a convincing talking face invites obvious misuse, and Google's answer is tiering. Customers can deploy from a curated library of pre-built avatars with no special permission. Generating a custom avatar โ from a single high-quality reference image and an audio sample, preserving likeness and brand styling โ sits behind a strict enterprise allowlisting and verification process.
Everything the feature outputs is watermarked with SynthID, woven into both the audio and the video stream rather than stamped on a corner of the frame. Google's stated reason is keeping synthetic content detectable to limit misattribution.
Where it runs, and what is still gated
The Google Cloud availability notice lists US and EU endpoints, provisioned throughput, enterprise compliance and data governance controls. The technology was first previewed at Google Cloud Next 2026. Gemini 3.8 Live Extended Thinking, the higher-reasoning sibling, remains in private preview.
What early customers report
Cox Automotive built a conversational avatar for Autotrader that highlights parts of the screen and calls tools to steer shoppers through search, comparison and financing. Marianne Johnson, the company's EVP and chief product officer, described the goal as letting buyers describe what they want in their own words instead of working through filters and menus.
Equal AI, which says it already fields more than a million live calls a day across nine Indian languages, credited the 3.8 Live release with better interruption handling and more reliable tool calls. Salesforce is pairing the model with Agentforce, and voice platform Specs singled out improvements to voice activity detection and overall latency.
What none of those statements include is a latency figure, a per-minute price, or a completion-rate comparison against the same workflow run without a face. Google has not published any of the three, which leaves buyers evaluating a rendering feature on demo footage rather than measured outcomes. Provisioned throughput being a prerequisite for serious deployment also suggests the video generation is expensive enough that Google would rather sell reserved capacity than metered calls.
Outlook
Live Avatar is an enterprise product, not a consumer Gemini feature, and the customer roster reads like contact-centre and intake work rather than general assistants. That is a narrower bet than it first appears, but it lands in the same contested space as Google's earlier low-latency voice models and rival full-duplex releases. The differentiator Google is testing is whether a face measurably improves completion rates on tasks users currently abandon. Google DeepMind has not published numbers on that yet.
FAQ
Is Gemini 3.8 Live with Live Avatar available to consumers?
No. The feature is generally available only in Gemini Enterprise, Google's business tier, and access runs through the Gemini Live API. Consumer Gemini apps are unaffected by this release.
Can a company make an avatar that looks like a specific real person?
Custom avatars can be generated from a single reference image, but Google gates that capability behind enterprise allowlisting and a verification process. All generated audio and video also carries a SynthID watermark so the output stays identifiable as AI-generated.
How many languages does the avatar support?
Gemini 3.8 Live understands and speaks 97 languages with automatic language detection, and Live Avatar adapts lip-sync and expressions across all of them. Google says the avatar can switch languages mid-conversation without degrading video fidelity.






