Google Unveils Most Advanced Voice Processing Models Yet
Google has announced the launch of Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its most advanced voice processing models yet. These new models are designed to address latency problems associated with voice-based AI agents by enabling near-real-time reasoning and simultaneous speech-and-thought processing.
The Gemini 3.8 Live model boasts top-tier benchmark results, achieving a score of 82.6 on the Artificial Analysis Speech to Speech Quality Index, surpassing GPT-Live-1-Astra and Grok Voice Think Fast 2.0. The Extended Thinking model also achieves impressive scores on various benchmarks, including T-Voice (68.6%), T-Voice-banking (35.1%), and Big Bench Audio (97.7%).
The new models support automatic language detection, can switch languages in mid-conversation, understand and generate speech in 97 languages, and use early verbal cues to acknowledge user prompts in a more natural way.
Developers will be able to integrate the models within their applications through Google partner platforms such as Vercel, Agora, LiveKit, Pipecat, Fishjam, and Vision Agents. The standard Gemini 3.8 Live model costs $0.005 per minute for audio inputs and $0.018 per minute for outputs.