Google Unveils High-Precision Speech-to-Text Model, Outpaces OpenAI Competitors
Google has launched a high-precision speech-to-text model called Gemini 3.5 Transcribe, which outperforms OpenAI's equivalent products. The model is designed to handle noise and specialized terminology, with applications in real-time captioning, voice agents, and post-call analysis.
Gemini 3.5 Transcribe offers both real-time and offline processing modes, with performance metrics that surpass competing products. It achieves a word error rate of 4% in streaming scenarios and 2.6% in non-streaming scenarios on the FLEURS benchmark.
The model is available to developers and select users in public preview, with general users able to access it via the Gemini application on macOS and Android. Google has stated that the model will soon arrive on the Chrome browser.