Google Unveils Speech-to-Text Model with Advanced Contextual Understanding
Google has released its latest speech-to-text model, Gemini 3.5 Transcribe, which boasts higher recognition accuracy and extends its capabilities to comprehend speakers' intentions and automatically organize output results.
This marks a significant step forward for speech recognition technology, enabling it to move beyond verbatim transcription and understand the context of spoken language.
Gemini 3.5 Transcribe can identify self-corrections, delete filler words such as 'um' and 'uh', and convert unstructured spoken language into formatted text.
The model supports over 85 languages, has automatic language recognition capability, and can handle scenarios where multiple languages are mixed.