Microsoft Introduces Real-Time Transcription for Voice AI Applications
Microsoft has introduced a new live transcription model, MAI-Transcribe-2-Streaming, as part of its AI lineup. Released on October 1, 2026, this model is paired with two speech generators, MAI-Voice-2.1 and its Flash variant, designed for voice agents that listen, respond, and speak in real-time. The system is positioned for applications like customer-service calls, captions, and voice assistants, reflecting a shift towards conversational AI.
The streaming transcription model provides provisional words as a person is speaking, allowing automated agents to begin processing requests before the speaker finishes. Microsoft claims support for 60 languages and automatic language detection, with initial transcript hypotheses appearing around 100 milliseconds after audio input. However, the model is still in a public preview phase, lacking a service-level agreement and not recommended for production workloads.
Pricing for the streaming service is set at $0.54 per hour of audio through the end of 2026, significantly higher than the $0.10 per hour for batch transcription. Microsoft also highlights its leadership on an Artificial Analysis streaming-speech leaderboard, with a 2.50% final word-error rate, outperforming competitors like xAI’s Grok Voice Transcribe 2.0. However, real-world performance may vary depending on accents, noise levels, and specific use cases.
For East African developers, particularly in Kenya, there are considerations around language support. While Swahili is listed among the base model’s 60 languages, the streaming version does not provide a detailed language-by-language table, and the voice models do not include Swahili. Builders are advised to test local accents, network conditions, and language coverage before deployment.