Google Unveils Most Powerful Speech-to-Text Model Yet
Google has released its most powerful speech-to-text model yet, Gemini 3.5 Transcribe, which can automatically process filler words, verbal slips, and repeated expressions during transcription.
The model is designed to convert original speech into text closer to the finished draft, reducing users' subsequent workload of organizing meeting minutes, interview shorthand, and call content.
Gemini 3.5 Transcribe supports over 85 languages and regional variants, can automatically determine the language currently in use, and can continue to transcribe even if the language is switched in one sentence or the same paragraph of dialogue.
The model also provides custom vocabulary lists, which developers can use to guide the model to prioritize recognizing specific contents. Google stated that this feature is usually more effective when the custom vocabulary list is controlled within 100 words.