Skip to content
Back to Guavy Wire
Stocks

Microsoft Unveils Trio of AI Speech Models with Advanced Language Support

Instruments
MSFT
Share

Microsoft has unveiled three new AI speech models designed for voice applications: MAI-Transcribe-2-Streaming, MAI-Voice-2.1, and its Flash variant.

The transcription model, MAI-Transcribe-2-Streaming, supports 60 languages with real-time, incremental speech-to-text and automatic language detection.

According to Artificial Analysis, this model ranked first for both final and partial transcript accuracy in its streaming evaluation.

Microsoft reports that the model can begin producing transcription just over 100 milliseconds after receiving audio.

The company also introduced MAI-Voice-2.1 and its Flash variant for text-to-speech, which supports 23 languages and generates 45 seconds of audio with about 150 milliseconds of end-to-end latency.

More on Stocks

Disclaimer: Guavy is a data and market intelligence provider, not an investment adviser. The information, signals, and market analysis provided by the Guavy API and related services are for informational purposes only and are not intended as financial advice, investment recommendations, or an endorsement of any particular trading strategy. Trading in volatile markets, including cryptocurrency, carries significant risk and may not be suitable for all investors. Past performance is not indicative of future results. Users should consult with a qualified financial professional before making any investment decisions. Guavy makes no guarantee of trading profits or financial returns.

Market sentiment intelligence for apps, funds & agents

Location

729 55 Ave SW
Calgary AB T2V 0G4
Canada

© 2026 Guavy Inc