Google Unveils Gemini 3.5 Transcribe with Real-Time Speech-to-Text in 85 Languages
Google has unveiled Gemini 3.5 Transcribe, an advanced speech-to-text model capable of transcribing real-time conversations in over 85 languages. This AI-powered tool not only recognizes spoken words but also auto-corrects verbal stumbles and filler words such as 'um'. It can process audio from streaming sources with a word error rate of 4.0 percent and recorded audio at 2.6 percent.
The model has been designed to work seamlessly in real-time, boasting 70 percent lower latency than its predecessor, Chirp 3. Gemini 3.5 Transcribe can be integrated into various platforms, including Google AI Studio, the Gemini Enterprise Agent Platform, and even the Gboard for Android via 'Rambler'. Additionally, support is coming soon for Chrome.
The model's capabilities extend beyond transcription, allowing it to hand off tasks such as image generation or web searches to other Gemini models through 'function calling'. This integration enables more complex interactions with users. The two interfaces provided by the model are the Live API and the Interactions API, catering to different use cases.