Microsoft Unveils Trio of AI Speech Models with Advanced Language Support
Microsoft has unveiled three new AI speech models designed for voice applications: MAI-Transcribe-2-Streaming, MAI-Voice-2.1, and its Flash variant.
The transcription model, MAI-Transcribe-2-Streaming, supports 60 languages with real-time, incremental speech-to-text and automatic language detection.
According to Artificial Analysis, this model ranked first for both final and partial transcript accuracy in its streaming evaluation.
Microsoft reports that the model can begin producing transcription just over 100 milliseconds after receiving audio.
The company also introduced MAI-Voice-2.1 and its Flash variant for text-to-speech, which supports 23 languages and generates 45 seconds of audio with about 150 milliseconds of end-to-end latency.