Meta Surpasses OpenAI and Google in Real-Time Transcription with Muse Voice Transcribe
Meta's Superintelligence Labs has released Muse Voice Transcribe, a real-time speech recognition model that outperforms its competitors in some benchmarks. The model can recognize over 20 speakers and has been trained on more than 70 languages, including cases where multilingual speakers switch languages during conversations.
Muse Voice Transcribe is available through the Meta Model API, Meta AI for Mac, and in Muse Code, with a pricing plan of $3.00 per 1,000 audio minutes or $0.18 per hour. Unlike previous models from Meta, the weights of this model will not be made open.
On Artificial Analysis's AA-WER Streaming speech-to-text accuracy benchmark, the model achieved a word error rate of 3.1%, outperforming competitors like Cartesia Ink-2 (3.4%) and ElevenLabs' Scribe v2 Real-time (3.6%). When it comes to recognizing distinct speakers, Muse Voice Transcribe leads the pack with a 17.5% error rate across several standard benchmarks.