Google Unveils Gemini 3.5 Transcribe with Improved Speech-to-Text Capabilities
Google has introduced Gemini 3.5 Transcribe, its latest speech-to-text model designed for precise and intelligent real-time transcription. This model is an improvement over its previous version, Chirp 3, offering better latency, accuracy, and language support.
Gemini 3.5 Transcribe can handle background noise, complex jargon, and disfluency cleanup, making it suitable for various applications such as voice agents, real-time captioning tools, and post-call analytics pipelines.
The model is available through two separate APIs: Real-time streaming and Pre-recorded audio processing. It achieves an average Word Error Rate (WER) of 4.0% for streaming and 2.6% for non-streaming use-cases, as measured by Artificial Analysis.