Geo News
Community curated by people like you
LatestAICryptoHealthWorld AffairsUS Politics
Microsoft releases MAI-Transcribe-2-Streaming real-time speech model
00

Microsoft releases MAI-Transcribe-2-Streaming real-time speech model

Oct 2, 2026

Microsoft AI has released MAI-Transcribe-2-Streaming, its first real-time speech-to-text model, alongside two new text-to-speech models. Evaluated as the top-performing model for accuracy by Artificial Analysis, the streaming model transcribes 60 languages and achieves a 2.5% Word Error Rate. This release marks a strategic push by Microsoft AI Chief Executive Mustafa Suleyman to reduce reliance on external partners like OpenAI and Anthropic by building in-house voice agent capabilities.

Features of MAI-Transcribe-2-Streaming

  • ▪Microsoft AI released MAI-Transcribe-2-Streaming, its first real-time streaming speech-to-text model, on October 1, 2026
  • ▪MAI-Transcribe-2-Streaming supports more than 60 languages and features automatic, continuous language detection
  • ▪MAI-Transcribe-2-Streaming accepts human speech via a WebSocket, continuously updating a transcript and confirming it as final once the speaker finishes

Performance of MAI-Transcribe-2-Streaming

  • ▪MAI-Transcribe-2-Streaming delivers its first partial transcript hypotheses within 320 milliseconds on average according to Microsoft, or just over 100 milliseconds according to Artificial Analysis
  • ▪MAI-Transcribe-2-Streaming achieved a 2.5% Word Error Rate at 0.13 seconds after the end of speech for final transcripts in Artificial Analysis testing
  • ▪Artificial Analysis ranked MAI-Transcribe-2-Streaming first out of 38 models for both final and first partial transcript accuracy

Integration and availability of MAI-Transcribe-2-Streaming

  • ▪Developers can integrate MAI-Transcribe-2-Streaming via the Realtime API for OpenAI-compatible WebSockets or the Azure Speech SDK
  • ▪MAI-Transcribe-2-Streaming is available on Microsoft's Vercel AI Gateway, Azure Voice Live, Microsoft Foundry, and the MAI Playground

Features of MAI-voice-2.1 models

  • ▪Both MAI-Voice-2.1 and MAI-Voice-2.1-Flash models can clone a voice from a few seconds of reference audio, with built-in safeguards to prevent misuse
  • ▪MAI-Voice-2.1 supports 23 languages and 26 locales, allowing it to speak multiple languages in the same voice with a native accent
  • ▪Microsoft AI released two text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash, alongside the streaming transcription model, MAI-Transcribe-2-Streaming, on October 1, 2026
  • ▪MAI-Voice-2.1-Flash is designed for lower latency and cost, hitting a latency of 150 milliseconds and generating 45 seconds of audio

Pricing of the new models

  • ▪MAI-Transcribe-2-Streaming's introductory rate of $0.54 per audio hour through the end of 2026 is higher than streaming rates from xAI's Grok Voice Transcribe 2.0 and Meta's Muse Voice Transcribe, compared to $0.10 per hour for the batch MAI-Transcribe-2 model
  • ▪MAI-Voice-2.1 is priced at $22 per million characters, while the MAI-Voice-2.1-Flash variant is priced at $15 per million characters

Strategic goals for the MAI family

  • ▪Microsoft AI aims to use the MAI model family, including MAI-Transcribe-2-Streaming, to power its Copilot agents in platforms such as Excel and Outlook
  • ▪Microsoft AI Chief Executive Mustafa Suleyman instructed researchers to focus on the MAI model family, including MAI-Transcribe-2-Streaming, to reduce and ultimately eliminate costs paid to Anthropic and OpenAI

Debatable claims

  • ▪Real-time streaming AI is too resource-intensive to justify its minor latency improvements
  • ▪AI transcription models are accurate enough to fully replace human transcribers
  • ▪The commercial benefits of voice-cloning technology justify its potential for misuse

3 sources

MarkTechPost
Microsoft AI Releases MAI-Transcribe-2-Streaming: #1 Real-Time Speech-to-Text Model on Artificial Analysis
View source article
SiliconANGLE
Microsoft targets ultra-realistic voice agents with its first streaming transcription model - SiliconANGLE
View source article
The Decoder
Microsoft AI releases new transcription and text-to-speech models for voice agents
View source article

Featured stories

View more in AI research & benchmarks

OpenAI fires three safety researchers for allegedly sharing confidential information

Oct 2, 2026 · 4 sources

Accenture reports strong AI demand with 400 clients on advanced AI projects

Oct 1, 2026 · 5 sources

Google announces Gemini 4 Argon in limited release to cybersecurity partners

Sep 30, 2026 · 12 sources

OpenAI and Synopsys partner to develop AI model for chip design

Sep 30, 2026 · 3 sources

Story comments

Loading comments…

Topics

AI research & benchmarksAI tools & productsMicrosoftAI voice & speech tools

Featured stories

View more in AI research & benchmarks

OpenAI fires three safety researchers for allegedly sharing confidential information

Oct 2, 2026 · 4 sources

Accenture reports strong AI demand with 400 clients on advanced AI projects

Oct 1, 2026 · 5 sources

Google announces Gemini 4 Argon in limited release to cybersecurity partners

Sep 30, 2026 · 12 sources

OpenAI and Synopsys partner to develop AI model for chip design

Sep 30, 2026 · 3 sources