Google Unveils New AI Models for Multilingual Conversational Voice Agents
Google has released two new AI models, Gemini 3.8 Live and Gemini 3.5 Transcribe, to help developers build low-latency conversational voice agents with multilingual support. The Gemini 3.8 Live model can run API calls in the background while streaming audio responses without interruptions, analyzing live images and videos at up to 1 frame per second.
Gemini 3.8 Live supports 97 languages, including those with realistic accents and automatic mid-conversation language switching. Sessions last up to 15 minutes for audio-only and 2 minutes for audio and video, priced at $0.005 per minute for audio input and $0.018 per minute for audio output.
The Gemini 3.5 Transcribe model is designed for transcription and offers two API endpoints: the Live API for real-time streaming with sub-second-latency captions and the Interactions API for pre-recorded audio files up to 1 hour, including diarization and word timestamps.