Microsoft Expands MAI Lineup with Live Transcription and Two New Voice Models
Microsoft has expanded its Microsoft Automated Innovation (MAI) lineup by introducing live transcription and two voice models. The new additions are MAI-Transcribe-2-Streaming, a streaming speech-to-text model that supports 60 languages, and the MAI-Voice-2.1 and MAI-Voice-2.1-Flash text-to-speech models, which cover 23 languages and 26 locales.
The MAI-Transcribe-2-Streaming model is designed to provide real-time transcription, with the first transcript hypotheses appearing just over 100 milliseconds after audio begins arriving. The model also includes automatic language detection, eliminating the need for developers to select the spoken language before every session.
Microsoft notes that the streaming model costs $0.54 per audio hour through 2026, which is significantly higher than its non-streaming MAI-Transcribe-2 model that launched at $0.10 per hour. However, the company argues that the premium is justified for applications where a transcript can trigger an action before the speaker finishes speaking.
The two voice models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash, are designed to handle the response, with MAI-Voice-2.1 focusing on more expressive, higher-fidelity speech, while MAI-Voice-2.1-Flash is a faster, lower-cost variant.
Microsoft is billing the three models by different units: transcription by audio hour and speech generation by character count. Developers need to compare each rate against their service's expected usage.