Geo News
Community curated by people like you
LatestAICryptoHealthWorld AffairsUS Politics
Meta Superintelligence Labs releases Muse Voice Transcribe, a unified real-time transcription model
00

Meta Superintelligence Labs releases Muse Voice Transcribe, a unified real-time transcription model

Sep 1, 2026

Meta Superintelligence Labs has released Muse Voice Transcribe, its first real-time audio perception model. The model collapses automatic speech recognition, speaker diarization for over 20 speakers, and endpointing into a single autoregressive system. Trained on over 70 languages with reinforcement learning, it uses an adaptive delay policy to optimize accuracy and speed. The model is available via Meta's Model API for $3.00 per 1,000 minutes and powers dictation in Meta AI for Mac.

Muse Voice Transcribe unified model

  • ▪Muse Voice Transcribe is part of the broader Muse Spark family of models designed around natural conversation patterns.
  • ▪Muse Voice Transcribe natively supports audio inputs exceeding one hour and can handle more than 20 speakers.
  • ▪Muse Voice Transcribe collapses automatic speech recognition, speaker diarization, and endpointing into a single autoregressive model.
  • ▪Meta Superintelligence Labs released Muse Voice Transcribe on September 1, 2026, as its first real-time audio perception model.

Streaming ASR adaptive delay

  • ▪Muse Voice Transcribe processes audio in 80-millisecond chunks at 12.5 Hz, transforming each chunk into a single soft token.
  • ▪Meta Superintelligence Labs trained Muse Voice Transcribe using reinforcement learning to optimize a per-word adaptive delay policy that balances speed and accuracy.
  • ▪On the Artificial Analysis AA-WER Streaming benchmark, Muse Voice Transcribe recorded a 3.1% final-transcript word error rate at 0.16 seconds after the end of speech.

Speaker diarization tokens

  • ▪Muse Voice Transcribe performs speaker diarization by inserting special turn and speaker identification tokens directly into the single autoregressive text stream.
  • ▪Meta reported a 17.5% average diarization error rate for Muse Voice Transcribe across the AMI-IHM, AMI-SDM, and VoxConverse benchmarks.

Endpointing detection mechanism

  • ▪Muse Voice Transcribe handles endpointing by using special tokens to mark the start and end of speech, trained jointly with streaming speech-to-text.
  • ▪The endpointing mechanism in Muse Voice Transcribe detects when a speaker has finished a thought versus pausing mid-sentence to prevent chopped transcripts.

Multilingual code-switching support

  • ▪Muse Voice Transcribe natively supports code-switching, allowing it to transcribe sentences that mix multiple languages mid-clause.
  • ▪Muse Voice Transcribe was trained on over 70 languages, with 25 languages extensively verified and recommended at launch.

API pricing deployment options

  • ▪Muse Voice Transcribe is available as a hosted API on the Meta Model API priced at $3.00 per 1,000 audio minutes.
  • ▪Muse Voice Transcribe powers dictation features in the Meta AI desktop application for Mac and is integrated into the Muse Code coding agent.
  • ▪Meta Superintelligence Labs has not released open weights for Muse Voice Transcribe, meaning there is no self-hosted deployment path.

3 sources

Engadget
Meta's new AI transcription model can distinguish between multiple speakers and languages in real-time - Engadget
View source article
Marktechpost
Meta Superintelligence Labs Releases Muse Voice Transcribe: One Real-Time Model for Streaming ASR, Diarization, and Endpointing
View source article
Cryptobriefing
MSL rolls out Muse Voice Transcribe, a real-time audio model with speaker diarization
View source article

Featured stories

View more in AI voice & speech tools

Sarvam AI releases Saaras V4 speech model covering 22 Indian languages

Sep 26, 2026 · 2 sources

OpenAI revenue hits $70 billion annualized rate as ChatGPT reaches 1.2 billion weekly users

Sep 29, 2026 · 13 sources

Anthropic releases Claude Sonnet 5.5 with 30% speed and cost improvements ahead of planned IPO

Sep 28, 2026 · 6 sources

OpenAI and Synopsys partner to develop AI model for chip design

Sep 30, 2026 · 3 sources

Story comments

Loading comments…

Related Projects

Meta Superintelligence Lab

Topics

AI voice & speech toolsAI tools & productsAI startupsMultimodal models

Featured stories

View more in AI voice & speech tools

Sarvam AI releases Saaras V4 speech model covering 22 Indian languages

Sep 26, 2026 · 2 sources

OpenAI revenue hits $70 billion annualized rate as ChatGPT reaches 1.2 billion weekly users

Sep 29, 2026 · 13 sources

Anthropic releases Claude Sonnet 5.5 with 30% speed and cost improvements ahead of planned IPO

Sep 28, 2026 · 6 sources

OpenAI and Synopsys partner to develop AI model for chip design

Sep 30, 2026 · 3 sources