Gemini 3.8 Live learns to think and talk at the same time
DeepMind shipped Gemini 3.8 Live and 3.8 Live Extended Thinking, voice models that keep a conversation running while reasoning and tool calls happen in the background — here's what shipped and why the concurrency matters more than the benchmark scores.

DeepMind announced two new voice models on September 15, 2026: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. Both are built for real-time spoken dialogue, and both are rolling out today across the Gemini API, Google AI Studio, Gemini Enterprise, Search, and Workspace.
What shipped
Gemini 3.8 Live is the scale model — tuned for fluid, low-cost conversation with real-time visual grounding. It processes what it sees in near real time, so it can answer questions about a whiteboard, a screen, or a chessboard as the conversation happens. It also auto-detects and switches between 97 languages mid-conversation, which matters for anyone building a voice agent that can't assume a single language per session.
Gemini 3.8 Live Extended Thinking is the reasoning model, built for multi-step tasks that need more than a quick reply. DeepMind reports it taking the top spot on Artificial Analysis's Speech to Speech Quality Index (82.6), leading agentic task completion at 68.6% on τ-Voice and 35.1% on Sierra's τ-Voice-banking benchmark, and scoring 97.7% on Big Bench Audio. The base 3.8 Live model placed second in the Speech Agent Arena. Treat these as directional rather than final — they're self-reported and several of the underlying benchmarks are new enough that independent replication is thin — but the ordering (Extended Thinking ahead on reasoning-heavy tasks, base Live competitive on general preference) matches how DeepMind is positioning the two models.
The part worth paying attention to
The benchmark scores are the least interesting part of this release. The more consequential change is architectural: both models now reason and speak simultaneously instead of going silent while they work.
In practice, that means the model can execute a tool call or API request in the background — look something up, kick off a multi-step workflow, query a system — while continuing to talk. Extended Thinking in particular uses verbal filler ("let me check that...") to acknowledge a request immediately, then narrates progress as a multi-step task runs, rather than leaving the user staring at silence until the whole chain finishes.

This is the actual engineering problem voice agents have struggled with: conversational latency and task latency are not the same thing, and forcing them onto one clock means either the agent stalls mid-task or it fakes fluency by answering before it actually knows the answer. Decoupling the two — talk continues on its own clock, reasoning and tool execution run on theirs, and they resynchronize when the background work resolves — is a more honest way to build a voice agent than either extreme. DeepMind cites ServiceNow's EVA-Bench results as evidence the models hold this balance on complex workflows without sacrificing conversational quality, which is the right thing to be measuring if the claim is about production readiness rather than demo polish.
Where you can use it
3.8 Live is live now in the Gemini API, AI Studio, and Search Live, with a private preview in Gemini Enterprise. 3.8 Live Extended Thinking is in the API and AI Studio, in the Gemini app, and in Docs for Google AI Pro/Ultra subscribers and Gmail and Keep for all Google AI subscribers, with Gemini Enterprise and Workspace business rollout in private preview. The Live API is already integrated by Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents, with Salesforce, Genspark, and Lumeris named as early enterprise partners. All generated audio carries a SynthID watermark.
My take
If you're building voice agents in production, the concurrency model is the thing to test, not the leaderboard position. Background tool execution with live narration is exactly the failure mode that makes voice UIs feel broken today — the agent goes quiet, the user doesn't know if it heard them, and they either repeat themselves or hang up. Whether "let me check that..." actually holds up under real latency variance, or starts to feel like a canned stall tactic once you've heard it fifty times, is the kind of thing you only learn by wiring it into an actual workflow. I'd start there before trusting any of the benchmark deltas between the two models.