Google Unveils Gemini 3.5 Transcribe for Real-Time Speech-to-Text
Google has released Gemini 3.5 Transcribe, a speech-to-text model designed for real-time transcription. The model boasts improved accuracy and reduced latency compared to its predecessor.
Gemini 3.5 Transcribe converts audio into formatted text, handling background noise, technical terminology, and speaker corrections with ease. According to Artificial Analysis, the model achieves a 4.0% word error rate for streaming applications and 2.6% for non-streaming use cases.
The technology also improves time to final transcription by 70% over the previous Chirp 3 model. Google makes the model available through two application programming interfaces: the Live API and the Interactions API.