Microsoft Introduces Real-Time Transcription and Expanded Voice AI Models
Microsoft has expanded its Microsoft AI (MAI) portfolio with three new speech models: MAI-Transcribe-2-Streaming, MAI-Voice-2.1, and MAI-Voice-2.1-Flash. The October 1, 2026, releases introduce real-time transcription and multilingual speech synthesis capabilities.
MAI-Transcribe-2-Streaming delivers provisional text as someone speaks, updating results with additional context and finalizing the transcript once the speaker finishes. Microsoft claims the model provides partial results in just over 100 milliseconds and supports 60 languages with automatic language detection. The model is designed for applications such as call centers, voice assistants, and real-time note taking. Microsoft says its model ranked first in both partial and final transcript accuracy in Artificial Analysis’s streaming evaluation.
This release follows Microsoft’s earlier MAI transcription models, with MAI-Transcribe-2 introduced in September 2026, expanding language coverage and adding features like automatic language identification and speaker attribution. The new streaming model brings Microsoft into competition with other developers in the live transcription space.
Microsoft also introduced MAI-Voice-2.1, expanding its voice model family to 23 languages. The model allows a single speaker identity to carry across all supported languages with native accents. It includes features like emotional controls, code-switching, and voice cloning. MAI-Voice-2.1-Flash, a lower-latency variant, is designed for high-volume, latency-sensitive workloads such as conversational agents and call centers.