Meta Launches Real-Time Audio AI Model Muse Voice Transcribe
Meta has launched Muse Voice Transcribe, its first real-time audio AI model, capable of transcribing speech from over 20 speakers simultaneously while handling multiple languages in a single stream. Announced by the company's Superintelligence Labs on September 1, 2026, the system is designed for applications ranging from live dictation to long-form audio transcription and is now available to developers and through Meta's desktop applications. The model's core capability is real-time speech-to-text that does not require separate post-processing steps. It can generate transcriptions with minimal delay, identify different speakers in recordings lasting more than an hour, and seamlessly manage conversations where speakers switch between languages, a process known as code-switching. According to Meta, Muse Voice Transcribe is trained on audio from over 70 languages, with 25 validated at its launch. In a notable feature for the Indian market, it natively supports Hindi, Tamil, Telugu, Kannada, and Malayalam. A key innovation is its adaptive delay mechanism, which aims to balance accuracy and speed.