Google AI introduces two advanced TTS models: Gemini 3.8 Flash
Google AI has introduced two new, highly expressive Text-to-Speech (TTS) models: Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. These tools allow for the creation of audio content in over 100 languages, offering a choice of 2000+ pre-recorded voices.
Users gain expanded control over the dialogue, including the ability to manage line-by-line pacing and add natural vocal cues (such as simulating "mmm-m-m" sounds), which is critical for generating long, continuous, and high-quality audio.
The models are separated by task: Gemini 3.8 Flash TTS is designed for high-quality creative content (games, podcasts), where the creation of unique vocal characters with detailed control is required. Meanwhile, Gemini 3.8 Flash-Lite TTS is a scalable engine for mass dubbing and voice agents, which automatically adjusts tone and tempo in real-time with low costs.
Why it matters
- —Enhanced Speech Generation Quality: New models provide seamless and highly expressive audio suitable for professional use.
- —Model Specialization: Separation into Flash (creative) and Flash-Lite (scale) allows users to choose the right tool depending on the task—from unique characters to mass dubbing.
- —Advanced Control: The ability to manage dialogue line-by-line and add natural pauses significantly increases the realism of the created content.
Key facts
- Models support over 100 languages and offer a choice of 2000+ ready-made voices.
- Gemini 3.8 Flash TTS is ideal for creating unique characters in games and podcasts.
- Gemini 3.8 Flash-Lite TTS is optimized for scaling, mass dubbing, and real-time voice agent operation.
- Functionality includes line-by-line dialogue management and the addition of natural voice cues.
The full text is in the original source. Here we provide a brief summary and key facts.