Alphabet Unveils Two Gemini Voice Models for Custom Audio Experiences
Alphabet's latest foray into AI-generated speech has brought forth two Gemini Voice Models. The models, dubbed Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, are designed to provide developers with the tools needed to create custom audio experiences without professional editing resources or a large budget.
The primary difference between the two models lies in their emphasis rather than raw capability. Gemini 3.8 Flash TTS is built for expressive range and granular control, aimed at audiobooks, podcasts, video game characters, and interactive media. Users can describe a voice in natural language, specifying role, accent, and vocal characteristics, and the model will generate it from scratch.
Gemini 3.8 Flash-Lite TTS, on the other hand, targets scenarios where volume matters as much as quality: dubbing, audio content production, and real-time voice agents. It retains fine-grained control over tone, pacing, and expressiveness but without the full creative toolset of its larger sibling.
The models are available immediately through the Gemini API and Google AI Studio, with enterprise access via Gemini Enterprise expected to follow. The company has also lined up integration partners for the new TTS models, including Agora, LiveKit, Pipecat, and Vercel.