AI Newsway

Google's New Voice Models Talk While They Think - and Cost Less

Gemini 3.8 Live removes the dead air from voice agents, then undercuts the models it beats

|4 min read0
AI Summary
Google DeepMind shipped Gemini 3.8 Live and 3.8 Live Extended Thinking on September 15, 2026, speech-to-speech models that execute tool calls while the conversation keeps running. Extended Thinking leads Artificial Analysis' Speech to Speech Quality Index at 82.6 and costs $3.50 per hour of input audio, below GPT-Live-1 Astra's $5.83. Google is pairing a quality lead with a lower price, a combination that usually forces rivals to respond on cost rather than capability.
Google's voice assistant running on a Pixel handset, the consumer surface where Gemini 3.8 Live's speech-to-speech models now land
Google's voice assistant running on a Pixel handset, the consumer surface where Gemini 3.8 Live's speech-to-speech models now land

The tell that you are talking to a machine is the pause. Ask a voice agent to look something up and it goes silent, because the model cannot speak and call a tool at the same time. Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, which Google DeepMind shipped on September 15, are built to remove that pause.

Key takeaways

  • Both models execute tool calls in the background while the conversation continues, and auto-detect switches across 97 languages without restarting the session.
  • Extended Thinking holds the top spot on Artificial Analysis' Speech to Speech Quality Index at 82.6, with GPT-Live-1 Astra at 81.5 and Grok Voice Think Fast 2.0 at 81.3.
  • An hour of input audio costs $0.84 on the standard model and $3.50 on Extended Thinking, against $5.83 for Astra.

What background execution actually buys

Extended Thinking reasons and speaks simultaneously. Instead of the silence, it offers an early verbal cue - something on the order of "let me check that" - then narrates its way through a multi-step job as the steps complete. The heavier model is the one aimed at that kind of work; the standard Gemini 3.8 Live is the cheaper, higher-volume sibling tuned for fluid dialogue and visual grounding.

Visual context is live in both. A user can hold a camera on a screen and ask about what is there, and the model folds that into the running exchange. Function calling happens off to the side, which is the difference between an assistant and the phone tree most AI agents still resemble.

How it compares against Astra and Grok

Against OpenAI's GPT-Live-1 Astra, the margins are thin on quality and wide on completion. Extended Thinking posts 68.6% on the tau-Voice agentic benchmark where Astra manages 67.9%, and on Sierra's tau-Voice-banking evaluation - which asks whether the agent resolves a customer-service request rather than merely sounding competent - it reaches 35.1% against Astra's 32.0%.

xAI's entries trail further back. Grok Voice Think Fast 2.0 scores 56.5% on tau-Voice, and xAI-Realtime manages 16.5% on the banking test. Google reports 97.7% for Extended Thinking on Big Bench Audio, while the standard model takes 76.0 on the speech index and second place in the Speech Agent Arena, a preference-based ranking rather than a scored one.

The price is the argument

Voice deployments are metered by the hour, so the cost table carries more weight than the leaderboard. Reading the Artificial Analysis figures, Grok Voice Think Fast 2.0 bills $4.80 per hour of input audio and Astra $5.83 - both above the $3.50 Google charges for the model that outscores them, and several times the $0.84 standard tier.

Undercutting while winning is a repeat move. The company priced Gemini 3.7 Flash at half its predecessor, and 3.8 Flash landed only two weeks before this release, so the audio line is now on the same cadence as the text one.

Who can use it

Developers reach both models through the Gemini API and Google AI Studio; enterprises get a private preview on Gemini Enterprise. LiveKit, Pipecat, Agora, LangChain and Vercel already surface the Live API, and Salesforce, Genspark and Lumeris are cited as early adopters.

Consumers meet Extended Thinking inside Gemini Live, Search Live and the Workspace apps - Docs, Gmail and Keep - on paid Google AI plans. Every second of generated audio carries a SynthID watermark, Google's audio provenance marker.

What none of this settles is whether the models hold up off-script. Evaluations run to a plan; a support queue does not, and the failure mode for Google DeepMind here is an agent that narrates confidently through a task it never completed. The price advantage buys room to find out in production.

FAQ

Is Gemini 3.8 Live available to developers today?

Yes. Both Gemini 3.8 Live and 3.8 Live Extended Thinking became accessible through the Gemini API and Google AI Studio on September 15, 2026. Enterprise access is a private preview on Gemini Enterprise, with Gemini Enterprise for Customer Experience listed as coming soon.

How much does Gemini 3.8 Live cost compared to GPT-Live-1 Astra?

On a cost-per-hour-of-input-audio basis using the Big Bench Audio subset, Gemini 3.8 Live is $0.84 and Extended Thinking is $3.50. GPT-Live-1 Astra is $5.83 an hour, putting Google's cheaper tier at roughly one-seventh the price.

How many languages do the models handle?

Ninety-seven, with automatic detection when a speaker switches mid-conversation. There is no need to restart the session or declare a language in advance for the change to register.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

A 25-Turn Pressure Test Shows LLMs Fold While Their Reasoning Holds
AI & Machine Learning

A 25-Turn Pressure Test Shows LLMs Fold While Their Reasoning Holds

A new arXiv benchmark called SPINE argues with models for up to 25 turns and finds collapse rates rise with conversation length for all seven systems tested.

Seung Jung3 days ago
An AI Cracked a 373-Year-Old Cipher. Then Someone Checked the Microfilm.
AI & Machine Learning

An AI Cracked a 373-Year-Old Cipher. Then Someone Checked the Microfilm.

Vals AI reported Claude Fable 5.1 solved a 373-year-old cipher in 44 minutes. An independent replication reports 8 of 64 letters match โ€” chance level.

Seung Jung3 days ago
100 DeepMind Agents Split Into Cheaters and Whistleblowers
AI & Machine Learning

100 DeepMind Agents Split Into Cheaters and Whistleblowers

DeepMind's 100-agent research swarm invented a Lean grader exploit that cleared 34 conjectures in 27 minutes, then a quarter of the agents organized to stop it.

Seung Jung2 days ago
Google's Agent Swarm Hits 71% on Research-Level Math Proofs
AI & Machine Learning

Google's Agent Swarm Hits 71% on Research-Level Math Proofs

Google's Stellar Colosseum harness scores 71.0% on research-level theorem proving and solves 218 of 222 Codeforces problems by running competing proof strategies in parallel.

Seung Jung2 days ago
DeepMind Precomputed 9 Billion DNA Variants Into a 1-Petabyte Atlas
AI & Machine Learning

DeepMind Precomputed 9 Billion DNA Variants Into a 1-Petabyte Atlas

The bottleneck in reading the human genome has never been sequencing it โ€” it has been working out which of the roughly 9 billion possible single-letter changes...

Seung Jung3 days ago
Astra Tripled Claude's Vending-Bench Profit โ€” and Refused to Collude
AI & Machine Learning

Astra Tripled Claude's Vending-Bench Profit โ€” and Refused to Collude

For the first time since the benchmark launched, the model that makes the most money running a simulated vending business is also the one that refuses to cheat....

Seung Jung3 days ago