Google Requires Consent Before Cloning Voices with AI
Google has recently introduced two new text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. These models are designed for various applications such as gaming, audiobooks, podcasts, and interactive media.
The system allows users to direct every performance line by line with granular control over acting cues, pacing, dialect shifts, and backchanneling. Scripts can carry cues such as laughs, sighs, and gasps. The company claims the models hold voice quality across hours of continuous audio with minimal speaker drift.
The system also includes a consent verification process for voice replication. Before creating a clone of someone's voice, Google requires a 30-second audio sample of their voice or a voice they have the rights to use. The person whose voice it is must say 'yes' out loud on tape. This recording is then compared with the sample to ensure authenticity.
The feature will not be available in certain regions, including Illinois, Texas, the EEA, UK, Switzerland, and India. Google does not explain why this is the case, but it may be due to laws that treat a voice as personal property or biometric data used for identification purposes.