Google Challenges Speech API Providers with Gemini 3.5 Transcribe
Google has entered the voice API market with its Gemini 3.5 Transcribe platform, which includes two speech-to-text products: a standard transcription model and Gemini 3.5 Transcribe Live.
The standard model achieved a word error rate of 2.6% on Artificial Analysis's AA-WER benchmark, placing it fifth overall among competitors.
The Live variant offers real-time transcription, responding in just 0.40 seconds after speech ends - a key figure for voice assistants and live captioning applications.
Alphabet is directly challenging dedicated speech API providers by bundling competitive transcription into the Gemini platform, giving existing Google Cloud customers an incentive to consolidate their voice workloads.