Microsoft Unveils AI-Powered Audio Models for Conversational Voice Agents
Microsoft has released three AI-powered audio models designed to improve conversational voice agents. MAI-Transcribe-2-Streaming is a real-time speech-to-text model that ranks first in accuracy for both final and partial transcripts, according to Artificial Analysis.
The model supports 60 languages with automatic language detection and can produce partial hypotheses within 100 milliseconds of receiving audio. This allows an agent to begin reasoning or calling tools before the speaker finishes speaking, while live captions can appear during speech.
Microsoft's internal tests found that its transcription model produces words twice as fast as its closest competitor for real-time dictation and subtitling. The introductory price for MAI-Transcribe-2-Streaming is $0.54 per audio hour through the end of the year.
The other two models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash, bring Microsoft's text-to-speech system to 23 languages and 26 locales. The models can clone a voice across supported languages from just a few seconds of reference audio and include consent guardrails to prevent misuse.