Google's Gemini 3.5 Transcribe Revolutionizes Speech-to-Text with Unprecedented Accuracy
Google has introduced Gemini 3.5 Transcribe, a speech-to-text model that is more precise than its predecessors and already powers several first-party products.
This new model excels in areas where traditional speech recognition systems struggle, such as handling background noise, complex jargon, and disfluency cleanup.
Gemini 3.5 Transcribe can automatically detect and transcribe over 85 languages, adapt to custom vocabulary, and accurately attribute speech in pre-recorded audio with timestamps for up to three speakers.
The model has shown strong performance across noisy environments, achieving an average Word Error Rate (WER) of 4.0% for streaming and 2.6% for non-streaming use-cases.