Microsoft Unveils AI Models for Real-Time Transcription and Voice Applications
Microsoft has released three new artificial intelligence models aimed at transcription and voice applications.
The MAI-Transcribe-2-Streaming model is a real-time transcription tool that can turn live speech into text in over 60 languages, with automatic language detection. This model produces its first hypotheses within hundreds of milliseconds of receiving audio and refines them as more context arrives.
This technology has significant implications for industries such as customer service, where agents can begin identifying a caller's request before the sentence is complete. The model also enables live transcription experiences that surface words almost as quickly as they are spoken.
The MAI-Transcribe-2-Streaming model is available in 60 languages and costs $0.54 per hour on an introductory basis through the end of the year. It has been found to be the top transcription model for accuracy, with words appearing as early as 320 milliseconds after they are spoken.
The company has also released two new voice models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash, which prioritize expressive fidelity and balance natural speech with responsiveness and cost at scale respectively. These models support 23 languages and 26 locales with a single, consistent cross-language voice.