AI Newsway

Gemini's Live Avatar Holds Its Lip-Sync While Switching Among 97 Languages

Google made the animated persona generally available in Gemini Enterprise, with SynthID watermarks on every frame and custom faces locked behind an allowlist.

|5 min read0
AI Summary
Google made Gemini 3.8 Live with Live Avatar generally available in Gemini Enterprise on September 24, giving its speech-to-speech model an animated face that lip-syncs across 97 languages and switches between them mid-conversation. The feature adds asynchronous tool calling so dialogue continues while data loads, and live camera and screen understanding. Custom avatars require enterprise allowlisting, and all output carries SynthID watermarks. Early users include Cox Automotive, Equal AI and Salesforce.
A real-time rendered animated character, the kind of continuously generated face Gemini's Live Avatar now produces during enterprise voice calls.
A real-time rendered animated character, the kind of continuously generated face Gemini's Live Avatar now produces during enterprise voice calls.

Google made Gemini 3.8 Live with Live Avatar generally available in Gemini Enterprise on Thursday, giving its real-time conversational model an animated on-screen persona that lip-syncs and shifts expression as it talks. The feature couples the model's native speech-to-speech dialogue with low-latency streaming video generation, so the system listens, watches a camera feed and answers with a visible face rather than a waveform.

Key takeaways

  • Live Avatar reached general availability in Gemini Enterprise on September 24, one week after the underlying Gemini 3.8 Live model shipped.
  • The avatar maintains lip-sync and facial expression across 97 languages and can switch between them mid-conversation without what Google calls visual drift.
  • Every generated audio and video stream carries an imperceptible SynthID watermark, and building a custom avatar from a reference photo requires enterprise allowlisting.

What Live Avatar actually generates

In its launch post from the Gemini audio team, Google frames the feature as native coupling rather than a bolt-on renderer: video generation and speech are produced together, which is what allows precise lip-sync and turn-taking instead of a mouth animation chasing an audio track.

The more practical addition is asynchronous tool execution. The avatar can fire a tool call and pull data in the background while the conversation keeps running, so a caller is not left staring at a frozen face during a lookup. Google's demo walks through a hotel check-in where the dialogue continues uninterrupted while function calls resolve behind the scenes.

Live visual understanding rounds it out. The model processes camera feeds and screen shares alongside audio simultaneously, which makes it a genuinely multimodal agent rather than a voice bot with a portrait attached. Google is positioning it for web, mobile and interactive kiosks.

How Google handled the identity problem

A model that can render a convincing talking face invites obvious misuse, and Google's answer is tiering. Customers can deploy from a curated library of pre-built avatars with no special permission. Generating a custom avatar โ€” from a single high-quality reference image and an audio sample, preserving likeness and brand styling โ€” sits behind a strict enterprise allowlisting and verification process.

Everything the feature outputs is watermarked with SynthID, woven into both the audio and the video stream rather than stamped on a corner of the frame. Google's stated reason is keeping synthetic content detectable to limit misattribution.

Where it runs, and what is still gated

The Google Cloud availability notice lists US and EU endpoints, provisioned throughput, enterprise compliance and data governance controls. The technology was first previewed at Google Cloud Next 2026. Gemini 3.8 Live Extended Thinking, the higher-reasoning sibling, remains in private preview.

What early customers report

Cox Automotive built a conversational avatar for Autotrader that highlights parts of the screen and calls tools to steer shoppers through search, comparison and financing. Marianne Johnson, the company's EVP and chief product officer, described the goal as letting buyers describe what they want in their own words instead of working through filters and menus.

Equal AI, which says it already fields more than a million live calls a day across nine Indian languages, credited the 3.8 Live release with better interruption handling and more reliable tool calls. Salesforce is pairing the model with Agentforce, and voice platform Specs singled out improvements to voice activity detection and overall latency.

What none of those statements include is a latency figure, a per-minute price, or a completion-rate comparison against the same workflow run without a face. Google has not published any of the three, which leaves buyers evaluating a rendering feature on demo footage rather than measured outcomes. Provisioned throughput being a prerequisite for serious deployment also suggests the video generation is expensive enough that Google would rather sell reserved capacity than metered calls.

Outlook

Live Avatar is an enterprise product, not a consumer Gemini feature, and the customer roster reads like contact-centre and intake work rather than general assistants. That is a narrower bet than it first appears, but it lands in the same contested space as Google's earlier low-latency voice models and rival full-duplex releases. The differentiator Google is testing is whether a face measurably improves completion rates on tasks users currently abandon. Google DeepMind has not published numbers on that yet.

FAQ

Is Gemini 3.8 Live with Live Avatar available to consumers?

No. The feature is generally available only in Gemini Enterprise, Google's business tier, and access runs through the Gemini Live API. Consumer Gemini apps are unaffected by this release.

Can a company make an avatar that looks like a specific real person?

Custom avatars can be generated from a single reference image, but Google gates that capability behind enterprise allowlisting and a verification process. All generated audio and video also carries a SynthID watermark so the output stays identifiable as AI-generated.

How many languages does the avatar support?

Gemini 3.8 Live understands and speaks 97 languages with automatic language detection, and Live Avatar adapts lip-sync and expressions across all of them. Google says the avatar can switch languages mid-conversation without degrading video fidelity.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

Google's Gemini Omni 1.1 Flash Makes Cheap Drafts the Point
LLM & Chatbots

Google's Gemini Omni 1.1 Flash Makes Cheap Drafts the Point

Gemini Omni 1.1 Flash adds 40-second scene extension, first and last frame control, video references and 4K upscaling, plus 360p drafts at a third the cost.

Seung Jung28 days ago
Gemini Lands on Windows With an Alt+Space Shortcut and a Bid for Your Desktop
LLM & Chatbots

Gemini Lands on Windows With an Alt+Space Shortcut and a Bid for Your Desktop

Google's native Gemini app for Windows opens over any application with Alt+Space, connects to Gmail and Drive, and is available globally on Windows 10 and 11.

Seung Jung14 days ago
GLM-5.3 Scores 60 on Artificial Analysis Index at Half the Usual Output Price
LLM & Chatbots

GLM-5.3 Scores 60 on Artificial Analysis Index at Half the Usual Output Price

Artificial Analysis scored Z.ai's GLM-5.3 at 60 on its Intelligence Index, well above the 35 median, at $4.40 per million output tokens. The catch is verbosity.

Seung Jung37 days ago
Grok 4.6 Reaches the AI Frontier Without Raising Its Price
LLM & Chatbots

Grok 4.6 Reaches the AI Frontier Without Raising Its Price

xAI's Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, matching GPT-5.6 Sol while holding pricing flat at $2/$6 per million tokens.

Seung Jung42 days ago
AWS Becomes the First Cloud to Carry OpenAI's Gated Cyber Models
SaaS & Cloud

AWS Becomes the First Cloud to Carry OpenAI's Gated Cyber Models

Daybreak Red and Blue are now sold through Amazon Bedrock, moving OpenAI's gated cyber models into enterprise cloud procurement and AWS governance.

Seung Jung41 days ago
Grok Bot Lets AI Agents Run Their Own Group Chat โ€” and Sign Into Your Accounts
LLM & Chatbots

Grok Bot Lets AI Agents Run Their Own Group Chat โ€” and Sign Into Your Accounts

SpaceXAI opened a Grok Bot beta where multiple agents coordinate in group chats, assign ownership to each other, and sign into a user's own accounts.

Seung Jung41 days ago