Microsoft Unveils Trio of Advanced Voice AI Models
Microsoft has introduced three new models for its artificial intelligence (AI) portfolio: MAI-Transcribe-2-Streaming, MAI-Voice-2.1, and MAI-Voice-2.1-Flash. These models are designed to help developers build faster, more natural, and multilingual conversational AI experiences.
The first model, MAI-Transcribe-2-Streaming, is a real-time speech-to-text model that can deliver low-latency transcription across 60 languages. It automatically detects the language being spoken and provides partial transcription results shortly after receiving the audio, with its first hypotheses appearing in just over 100 milliseconds.
The other two models focus on multilingual text-to-speech capabilities: MAI-Voice-2.1 supports 23 languages and 26 locales, while MAI-Voice-2.1-Flash is a faster version aimed at high-volume and latency-sensitive applications. The Flash model can generate 45 seconds of audio with end-to-end latency of around 150 milliseconds.
The new models are available through Microsoft's AI services, including the public preview for MAI-Transcribe-2-Streaming, and are priced at $22 per 1 million characters for MAI-Voice-2.1 and $15 per 1 million characters for MAI-Voice-2.1-Flash.