Gemini 3.8 Flash TTS & Flash-Lite TTS Voice Generation Models
Published by Google on its official blog on 23 September 2026 and described as its most expressive audio generation models yet: Gemini 3.8 Flash TTS is built for deep creative direction and character design, creating entirely new voices from natural language and directing every performance line by line, while Gemini 3.8 Flash-Lite TTS is built for high-volume, cost-efficient dubbing; both offer 2,000+ production-ready voices across more than 100 languages and dialects and voice replication from just a 30-second audio sample, backed by built-in consent verification, SynthID watermarking and C2PA credentials
Tool Overview
Features, steps and FAQ below
Features
- ✓ Two distinct models: Google says Gemini 3.8 Flash TTS is built for deep creative direction and character design, creating entirely new voices from natural-language prompts and directing every line with granular control over acting cues, pacing, dialect shifts and backchanneling, while Gemini 3.8 Flash-Lite TTS is built for high-volume, cost-efficient dubbing and expressive voice agents
- ✓ Generative voice design and a large library: Google says Gemini 3.8 Flash TTS customizes role, accent and voice characteristics across more than 100 languages and dialects using natural-language prompting, and offers access to 2,000+ production-ready voices including regional varieties like Mexican Spanish, Quebec French and Scots English
- ✓ Voice replication and safety: Google says a consistent vocal profile can be recreated from just a 30-second audio sample of your voice or a voice you have the rights to use, backed by built-in consent verification (a verbal consent recording from the voice owner matching the reference speaker), SynthID watermarking and C2PA credentials; voice replication through AI Studio is not available in Illinois, Texas, the EEA, the UK, Switzerland and India
- ✓ Line-by-line direction and long-form audio: Google says you can write stage directions or let Gemini steer delivery with natural script cues, maintain high voice quality with minimal speaker drift across hours of continuous audio, and use native two-speaker scene staging plus scripted vocal bursts and backchanneling
- ✓ Benchmarks and availability: Google says Gemini 3.8 Flash TTS secures the #1 overall spot on Hume AI's Voice Design Benchmark (71.4) and leads accent modeling (60.8), with Gemini 3.8 Flash TTS and Flash-Lite TTS taking the #1 and #2 spots on Hume AI's Overall Quality Index; it is rolling out in the Gemini API, Google AI Studio and Gemini Notebook, with Flash-Lite TTS also in Google Vids
How to Use
- Developers can use Gemini 3.8 Flash TTS and Flash-Lite TTS directly in the Gemini API and Google AI Studio speech generation
- In Google AI Studio's voice design workspace, prompt entirely new voices from scratch or replicate a voice you have the rights to use
- Bring the designed voices into the dual-speaker screenplay editor to direct emotion, pacing and dialect line by line and add backchanneling and vocal detail
- Save the voices for consistent performance across projects and use the output for audiobooks, podcasts, games or real-time voice agents
FAQ
What is Gemini 3.8 Flash TTS?
Google says Gemini 3.8 Flash TTS is a new text-to-speech model in the Gemini family built for deep creative direction and character design, creating new voices from natural language and directing every line, with 2,000+ voices across more than 100 languages.
How is it different from Flash-Lite TTS?
Google says Gemini 3.8 Flash TTS is built for deep creative direction and character design such as gaming, immersive audiobooks and podcasts, while Gemini 3.8 Flash-Lite TTS is built for high-volume, cost-efficient dubbing, audio content creation and expressive voice agents.
How does voice replication work?
Google says voice replication recreates consistent vocal profiles from just a 30-second audio sample, requires a verbal consent recording from the voice owner that matches the reference speaker, and is backed by built-in consent verification, SynthID watermarking and C2PA credentials; voice replication through AI Studio is not available in Illinois, Texas, the EEA, the UK, Switzerland and India.
What safeguards protect the generated audio?
Google says every audio clip generated by its Gemini Audio models is watermarked with SynthID, an imperceptible watermark that keeps AI-generated speech detectable to help prevent misinformation, and voice replication requires consent verification and C2PA credentials.
Where is it available?
Google says Gemini 3.8 Flash TTS is rolling out in the Gemini API, Google AI Studio and Gemini Notebook, with Gemini Enterprise coming soon via API, and Gemini 3.8 Flash-Lite TTS is rolling out in the Gemini API, Google AI Studio and Google Vids.