Microsoft Unveils AI-Powered Voice Agent Models for Seamless Conversations
Microsoft has released two new AI-powered models for voice agents: MAI-Transcribe-2-Streaming and two text-to-speech models. The transcription model, which ranks first in accuracy on Artificial Analysis, can transcribe 60 languages in real-time with partial results delivered in under 100 milliseconds.
This allows voice agents to respond while someone is still speaking, making conversations more seamless. The introductory price for an hour of audio is $0.54 until the end of the year. Microsoft also released two new text-to-speech models: MAI-Voice-2.1 and its variant, MAI-Voice-2.1-Flash.
These models can speak 23 languages with a native accent in each one and have a latency of 150 milliseconds for the Flash variant. The cost for the Flash variant is $15 per million characters, compared to $22 for regular use. Both voice models can clone a voice from just a few seconds of reference audio.
Microsoft has built-in safeguards to prevent misuse, but the potential for abuse remains a concern. In one test, about half of 4,000 participants thought the voices belonged to real people.