Google Unveils Advanced Text-to-Speech Models with Custom Voice Design
Google has introduced two new text-to-speech (TTS) models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, which allow users to design custom voices, replicate existing voices, and generate multilingual audio at scale.
The models support more than 100 languages and provide access to over 2,000 production-ready voices. The flagship model, Gemini 3.8 Flash TTS, is designed for applications requiring nuanced acting and detailed performance control, such as audiobooks, studio narration, games, and podcasts.
Gemini 3.8 Flash-Lite TTS is the faster, lower-cost option, intended for high-volume dubbing and audio production. Both models support prebuilt voices, custom voice design, and voice replication, with users able to create voices by describing their desired role, accent, and vocal characteristics in natural language.
The new release adds fine-grained performance controls, allowing users to control tone, emotion, pacing, and other aspects of individual lines. Voice replication requires a 10-30 second reference recording of the user's own voice or one they have the rights to use, along with a separate consent recording from the speaker.