Google's Gemini 3.5 Transcribe Revolutionizes Voice Input
Google has released Gemini 3.5 Transcribe, an advanced audio model that can automatically detect and transcribe speech from over 85 languages. Unlike traditional transcription systems, Gemini 3.5 Transcribe is designed to understand how people actually talk, rather than just transcribing words.
The model achieves this by stripping filler words, following self-corrections, and formatting the result into usable text. It also attributes speech to up to three speakers in pre-recorded audio with word-level timestamps, a feature that can be particularly useful for processing recorded meetings, interviews, or podcasts.
Gemini 3.5 Transcribe already powers dictation on Android and the Gemini app on macOS, but Chrome support is next. This will allow users to dictate into any field on any web page, making voice input a more practical option for everyday computing.