Google Unveils Advanced AI Models for Real-Time Audio Reasoning
Google has unveiled Gemini 3.8 Live and 3.8 Live Thinking AI models, designed for real-time audio reasoning and multi-step tasks. The new models are part of Google's Gemini series, which enables high-complexity tasks with advanced multi-step reasoning and increased intelligence.
Gemini 3.8 Live is a dedicated speech-to-text model that offers highly precise transcription across 85+ languages with low error rates. The average Word Error Rate (WER) for streaming transcription is 4.0%, while non-streaming transcription has an average WER of 2.6%.
The Gemini AI can be invoked via smartphones using voice commands to perform complex multi-step tasks across Google and third-party apps. Additionally, all audio generated by the new models is watermarked with SynthID to detect AI-generated content and prevent misinformation.