Meta Debuts Muse Voice Transcribe for Real-Time Multilingual Audio
The Meta Superintelligence Lab model supports 20-plus speakers and handles simultaneous language switching.
Meta Superintelligence Lab (MSI) has released Muse Voice Transcribe, the company's first real-time audio perception model. The tool introduces streaming automatic speech recognition and advanced diarization to the Meta ecosystem, marking a significant step in how the company processes live voice data.
The model distinguishes between more than 20 different speakers in real time. According to Meta Research, Muse Voice Transcribe ranks first on Artificial Analysis for streaming speech-to-text and public diarization benchmarks as of September 1, 2026. To ensure accuracy, the system employs an "adaptive delay" mechanism. Mark Zuckerberg explained that the model "decides when to listen," waiting longer to process difficult words while committing faster to easy ones to predict tokens more effectively.
Global Language Support
Beyond speaker identification, the model targets a global audience. It was trained across more than 70 languages, with 25 validated at launch. A key technical capability is its support for "code-switching," which allows the AI to transcribe sentences containing multiple languages simultaneously without losing coherence.
This release arrives during intense competition in the audio AI space, launching less than a week after Google introduced Gemini 3.5 Transcribe. Muse Voice Transcribe is part of a broader offensive from MSI, following the recent debuts of the Muse Glimmer open-weight model and the Muse Code coding agent.
Market Integration and Access
Meta is positioning the technology as a system-wide utility. The model is currently available through the Meta AI Mac app and Muse Code. For developers, Meta has opened the Model API with a pricing structure set at $3 per 1,000 audio minutes.
By integrating these capabilities into a desktop application and a scalable API, Meta is targeting critical needs in accessibility and global business communication. The ability to accurately track dozens of speakers in a multilingual environment provides a foundation for more sophisticated AI-driven productivity tools and real-time translation services.
Future Outlook
As Meta continues to roll out the Muse suite, the industry will be watching how these tools integrate with the company's broader hardware and software ecosystem. While current benchmarks place Muse Voice Transcribe at the top of the streaming speech-to-text category, the long-term challenge will be maintaining this accuracy across the full spectrum of 70-plus languages used during training.