Google launches Gemini 3.8 Live and 3.8 Live Extended Thinking: reasoning while speaking, with live switching across 97 languages

2026-09-18·10 min read

On September 15, 2026, Google published Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on its official blog, credited to Tom Ouyang, Principal Engineer, and Malini Jaganathan, Member of Technical Staff, on behalf of the Gemini Audio team. The split in positioning is explicit: 3.8 Live is built for scale and cost efficiency, combining conversational intelligence with fluid dialogue and visual grounding; 3.8 Live Extended Thinking is built for high-complexity tasks, aiming at increased intelligence and multi-step reasoning. Google framing in the update is that the two models give developers and enterprises the building blocks for reliable, production-ready voice agents. That last phrase carries the weight - what Google is selling is not merely nicer conversation, but the ability to drive a multi-step task to completion by voice.

The scores are the first thing worth recording. 3.8 Live Extended Thinking captures the number one overall spot on Artificial Analysis Speech to Speech Quality Index at 82.6; it scores 68.6% on tau-Voice and 35.1% on Sierra tau-Voice-banking, which Google uses to argue for leadership in agentic task completion. On reasoning it reaches 97.7% on Big Bench Audio while maintaining what Google calls a highly competitive price point against other frontier models. 3.8 Live, meanwhile, secures second place in the Speech Agent Arena on user preference, with Google stressing its cost effectiveness at scale. On ServiceNow EVA-Bench, a benchmark for evaluating voice agents, Google says both models push the Pareto frontier for complex workflows by balancing accuracy with conversational quality, and notes specifically that this run was done on the Live API on Gemini Enterprise Agent Platform. Note the benchmark names and the platform qualifier when citing these figures - they come from Google own blog.

The second thing worth recording is engineering detail. 3.8 Live processes visual inputs in near real-time, enriching conversations with context; it automatically detects and transitions between 97 supported languages mid-conversation, so a user switching language does not need to start a new session. Easier to overlook is background execution: 3.8 Live executes tools and API calls in the background while continuing the conversation, so it can acknowledge a request and keep talking while the task finishes behind it. 3.8 Live Extended Thinking takes a different route: it reasons and speaks simultaneously. Google describes the mechanics - early verbal cues such as Let me check that acknowledge prompts naturally, and live progress narration walks users through multi-step background tasks. The intent is clear: to remove the silent gaps that have made voice agents frustrating. The blog demos include guiding employee onboarding in real time, playing chess in near real-time, turning rough sketches plus live voice feedback into functional React components, coordinating multi-step bookings with asynchronous function calls, and generating complete business plans and custom marketing toolkits on the fly by speech.

Availability and ecosystem are the third block of information. For developers, both models are available in the Gemini API and Google AI Studio; for enterprises, they enter private preview in Gemini Enterprise, coming soon to Gemini Enterprise for Customer Experience and to Google Workspace business customers. For everyone else, 3.8 Live arrives in Search Live, and 3.8 Live Extended Thinking arrives in Gemini Live, for Google AI Pro and Ultra subscribers in Workspace Docs, and for all Google AI subscribers in Gmail and Keep. Google also lists developer platforms building on the Gemini Live API: Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel and Vision Agents, which handle complex real-time media streaming infrastructure so developers can focus on the user experience. The partner quotes come from Salesforce, Genspark and Lumeris, with Google saying they highlight latency, fluidity and tool-calling capability. One caveat matters: the blog says both models are rolling out starting today, and a phased rollout means not every user sees the entry point at the same time.

Finally, two things on safety and sourcing. Google states explicitly that all audio generated by its AI products is watermarked with SynthID, an imperceptible watermark woven directly into the audio output so that AI-generated content remains detectable and can help prevent misinformation, with details of its safety and responsibility approach in the model card. On sourcing: the post is dated September 15, 2026, and the page notes it was updated September 17, 2026, so check that gap before citing benchmark figures; the page also links two related posts, on the launch of Gemini 3.8 Flash and 3.8 Flash Cyber, and on proactive cyber defence for governments and enterprises. One more caution: hourly pricing figures circulating in search results do not appear in the body of this official post, so this article deliberately omits specific price numbers and keeps only the official cost-efficiency and competitiveness wording. If you want a citing rule: take scores and availability from the official blog, and take pricing from the official pricing page.

🤔 Frequently Asked Questions

Q1: What is the difference between Gemini 3.8 Live and 3.8 Live Extended Thinking?

Gemini 3.8 Live is built for scale and cost efficiency, combining conversational intelligence with fluid dialogue and visual grounding, and places second in the Speech Agent Arena. 3.8 Live Extended Thinking is built for high-complexity tasks, aiming at increased intelligence and multi-step reasoning, and takes the number one spot on Artificial Analysis Speech to Speech Quality Index at 82.6. Roughly: the former suits large-scale low-cost deployment, the latter suits complex workflows that need deeper reasoning.

Q2: Which benchmark results did Google publish?

For 3.8 Live Extended Thinking: number one overall on Artificial Analysis Speech to Speech Quality Index at 82.6, 68.6% on tau-Voice, 35.1% on Sierra tau-Voice-banking, and 97.7% on Big Bench Audio. Gemini 3.8 Live places second in the Speech Agent Arena. On ServiceNow EVA-Bench, Google describes both models as pushing the Pareto frontier for complex workflows, with that run performed on the Live API on Gemini Enterprise Agent Platform.

Q3: Where can regular users try them?

Gemini 3.8 Live arrives in Search Live. Gemini 3.8 Live Extended Thinking arrives in Gemini Live, for Google AI Pro and Ultra subscribers in Workspace Docs, and for all Google AI subscribers in Gmail and Keep. Developers get them in the Gemini API and Google AI Studio, enterprises in private preview on Gemini Enterprise. Google says both are rolling out starting the day of launch, so the entry point appears at different times for different users.

Q4: Is generated audio watermarked?

Yes. Google states that all audio generated by its AI products is watermarked with SynthID, an imperceptible watermark woven directly into the audio output so that AI-generated content stays detectable and can help prevent misinformation, with details of its safety and responsibility approach in the model card.

🛠️ Recommended Tools

  • Text to Speech - The fastest way to get a feel for live speech model quality is to generate your own audio from the same script and compare it against the official demos.
  • Speech to Text - Mid-conversation switching across 97 languages means transcripts hit language jumps. Running your own recordings through a transcription tool shows whether the switch points are recognised correctly.
  • Percentage Calculator - Metrics like 82.6, 68.6%, 35.1% and 97.7% use different scales. Putting them through a percentage lens keeps you from comparing scores across different benchmarks as if they were the same thing.

What stayed with me from this launch is not the scores but the fact that the phrase Let me check that made it into an official blog post. The worst moment with a voice assistant used to be the ten-plus seconds of silence after you finish a sentence, when you cannot tell whether it is thinking, stuck, or dead. Google answer is to have the model narrate its thinking - acknowledge you with a short verbal cue, then reason while reporting progress. That amounts to admitting the core problem in voice interaction is not recognition accuracy but silence. The other detail is switching between 97 languages mid-conversation: multilingual dialogue used to require picking a language up front or facing garbled recognition, and that boundary is now gone. As for executing tool calls in the background, it turns voice from a question-and-answer surface into a task surface, which is where the real shift in this generation sits. There are caveats too - one benchmark result carries an explicit platform note, and a rolling rollout means you and I may not see the entry point today. Check your own account to know for sure.

Summary

On September 15, 2026 (updated September 17), Google official blog introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, credited to Tom Ouyang and Malini Jaganathan on behalf of the Gemini Audio team. Gemini 3.8 Live is built for scale and cost efficiency, combining conversational intelligence, fluid dialogue and visual grounding, and places second in the Speech Agent Arena; Gemini 3.8 Live Extended Thinking is built for high-complexity tasks with increased intelligence and multi-step reasoning, and takes the number one spot on Artificial Analysis Speech to Speech Quality Index at 82.6, with 68.6% on tau-Voice, 35.1% on Sierra tau-Voice-banking and 97.7% on Big Bench Audio. Gemini 3.8 Live processes visual inputs in near real-time, automatically detects and transitions between 97 supported languages mid-conversation, and executes tools and API calls in the background while the conversation continues; 3.8 Live Extended Thinking reasons and speaks simultaneously, acknowledging prompts with early verbal cues such as Let me check that and narrating multi-step background tasks live. On ServiceNow EVA-Bench both models are described as pushing the Pareto frontier for complex workflows, with the run performed on the Live API on Gemini Enterprise Agent Platform. For developers both are live in the Gemini API and Google AI Studio; for enterprises in private preview in Gemini Enterprise, coming soon to Gemini Enterprise for Customer Experience and Google Workspace business customers; for everyone in Search Live, Gemini Live, Workspace Docs for Google AI Pro and Ultra subscribers, and Gmail and Keep for all Google AI subscribers. Ecosystem partners include Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel and Vision Agents, with Salesforce, Genspark and Lumeris quoted. All generated audio carries a SynthID watermark. Both models are rolling out starting launch day. All details come from Google official blog; no price figure absent from the official text is cited here. Primary source: Google official blog (published September 15, 2026, updated September 17, 2026).

Sources: Google Blog: Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking