Meta Superintelligence Labs has released Muse Voice Transcribe, its first real-time audio perception model. The model collapses automatic speech recognition, speaker diarization for over 20 speakers, and endpointing into a single autoregressive system. Trained on over 70 languages with reinforcement learning, it uses an adaptive delay policy to optimize accuracy and speed. The model is available via Meta's Model API for $3.00 per 1,000 minutes and powers dictation in Meta AI for Mac.
Sep 26, 2026 · 2 sources
Sep 29, 2026 · 13 sources
Sep 28, 2026 · 6 sources
Sep 30, 2026 · 3 sources
Story comments
Loading comments…