Microsoft AI Unveils World's Most Accurate Real-Time Audio Transcription Model
Microsoft AI has launched MAI-Transcribe-2-Streaming, a real-time audio transcription model that achieves the highest accuracy in the world. According to Microsoft AI CEO Mustafa Suleyman, the new voice architecture is 55% faster and 60% cheaper than comparable offerings like ElevenLabs.
The key feature of MAI-Transcribe-2-Streaming is its ability to provide continuous, low-latency live speech-to-text across 60 languages with built-in automatic language detection. This means that the model can output initial text hypotheses within roughly 100 milliseconds of receiving incoming audio and refine words as additional context arrives.
The model has secured the No. 1 spot for accuracy across both partial and completed transcripts on the independent benchmark platform Artificial Analysis, eliminating the historical tradeoff between high precision and sluggish processing times.