Gemini 3.8 Live & 3.8 Live Extended Thinking Real-Time Voice Models
Published by Google on its official blog on 15 September 2026 and described as its most advanced live dialogue models yet: Gemini 3.8 Live is built for scale and cost efficiency, combining conversational intelligence with fluid dialogue and visual grounding, automatically detecting and switching between 97 supported languages mid-conversation while executing tools and API calls in the background; 3.8 Live Extended Thinking is built for high-complexity tasks, reasoning and speaking simultaneously with live progress narration, capturing the #1 overall spot on Artificial Analysis' Speech to Speech Quality Index at 82.6
Tool Overview
Features, steps and FAQ below
Features
- ✓ Two model profiles: Google says Gemini 3.8 Live is built for scale and cost efficiency, combining conversational intelligence with fluid dialogue and visual grounding, while 3.8 Live Extended Thinking is built for high-complexity tasks with increased intelligence and multi-step reasoning
- ✓ Benchmark performance: Google says 3.8 Live Extended Thinking captures the #1 overall spot on Artificial Analysis' Speech to Speech Quality Index at 82.6, scores 68.6% on τ-Voice and 35.1% on Sierra's τ-Voice-banking, and 97.7% on Big Bench Audio, while 3.8 Live secures second place in the Speech Agent Arena and stays highly cost-effective
- ✓ Near real-time vision and multilingual support: Google says Gemini 3.8 Live processes visual inputs in near real-time to enrich conversations with context and automatically detects and transitions between 97 supported languages mid-conversation
- ✓ Background execution and thinking while speaking: Google says 3.8 Live executes tools and API calls in the background while continuing the conversation, and 3.8 Live Extended Thinking reasons and speaks simultaneously, using early verbal cues like "Let me check that…" and live progress narration for multi-step background tasks
- ✓ SynthID watermarking and developer ecosystem: Google says all generated audio is watermarked with SynthID, and the Gemini Live API plus platforms such as Agora, LangChain, LiveKit, Pipecat and Vercel let developers build and deploy high-performance voice interfaces
How to Use
- Developers can use 3.8 Live and 3.8 Live Extended Thinking in the Gemini API and Google AI Studio
- Enterprises can use them in private preview in Gemini Enterprise (with 3.8 Live Extended Thinking also coming to Google Workspace business customers)
- Everyone can use 3.8 Live in Search Live, and subscribers can use it in Gemini Live and in Workspace across Docs (Google AI Pro and Ultra) and Gmail and Keep (all Google AI subscribers)
- Developers can build and deploy voice-driven interfaces through the Gemini Live API and supported third-party platforms
FAQ
What are these two models?
Google says Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are its most advanced live dialogue models yet, with the former built for scale and cost efficiency and the latter built for high-complexity tasks requiring increased intelligence and multi-step reasoning.
How can you try them?
Google says 3.8 Live rolls out starting that day for developers in the Gemini API and Google AI Studio, for enterprises in private preview in Gemini Enterprise, and for everyone in Search Live; 3.8 Live Extended Thinking also rolls out in the Gemini API, AI Studio and Gemini Live and some Workspace subscription tiers.
How does it perform?
Google says 3.8 Live Extended Thinking captures the #1 overall spot on Artificial Analysis' Speech to Speech Quality Index at 82.6, scores 68.6% on τ-Voice, 35.1% on Sierra's τ-Voice-banking and 97.7% on Big Bench Audio.
How many languages are supported?
Google says Gemini 3.8 Live automatically detects and transitions between 97 supported languages mid-conversation.
Can the audio be detected as AI-generated?
Google says all audio generated by its AI products is watermarked with SynthID, an imperceptible watermark woven into the output to help detect AI-generated content and prevent misinformation.