Google Unveils Gemini 3.5 Transcribe: Real-Time Speech-to-Text Model
Google has released Gemini 3.5 Transcribe, a speech-to-text model that can convert raw audio into clean, formatted text in real-time. This tool is designed to automatically strip out filler words and self-corrections people make while speaking.
The new model supports over 85 languages with automatic detection and can handle multi-language code-switching. It also attributes speech to up to three speakers in pre-recorded audio, complete with word-level timestamps.
According to Google's developer documentation, the model is based on Gemini's broader audio understanding capabilities, with utterance-based language detection layered on top of the core transcription.
The company claims that Transcribe represents a major advancement over its previous transcription engine, Chirp 3. Time to final transcription improves by 70 percent, and the live-speech word error rate sits at 5.50 percent, down from Chirp 3's 7.32 percent.
Transcribe is available now in English for macOS Gemini app users and through the Rambler dictation feature on Android in select countries and languages, with Chrome support coming soon.