Microsoft AI has released MAI-Transcribe-2-Streaming, its first real-time speech-to-text model, alongside two new text-to-speech models. Evaluated as the top-performing model for accuracy by Artificial Analysis, the streaming model transcribes 60 languages and achieves a 2.5% Word Error Rate. This release marks a strategic push by Microsoft AI Chief Executive Mustafa Suleyman to reduce reliance on external partners like OpenAI and Anthropic by building in-house voice agent capabilities.
Features of MAI-Transcribe-2-Streaming
- ▪Microsoft AI released MAI-Transcribe-2-Streaming, its first real-time streaming speech-to-text model, on October 1, 2026
- ▪MAI-Transcribe-2-Streaming supports more than 60 languages and features automatic, continuous language detection
- ▪MAI-Transcribe-2-Streaming accepts human speech via a WebSocket, continuously updating a transcript and confirming it as final once the speaker finishes
Performance of MAI-Transcribe-2-Streaming
- ▪MAI-Transcribe-2-Streaming delivers its first partial transcript hypotheses within 320 milliseconds on average according to Microsoft, or just over 100 milliseconds according to Artificial Analysis
- ▪MAI-Transcribe-2-Streaming achieved a 2.5% Word Error Rate at 0.13 seconds after the end of speech for final transcripts in Artificial Analysis testing
- ▪Artificial Analysis ranked MAI-Transcribe-2-Streaming first out of 38 models for both final and first partial transcript accuracy
Integration and availability of MAI-Transcribe-2-Streaming
- ▪Developers can integrate MAI-Transcribe-2-Streaming via the Realtime API for OpenAI-compatible WebSockets or the Azure Speech SDK
- ▪MAI-Transcribe-2-Streaming is available on Microsoft's Vercel AI Gateway, Azure Voice Live, Microsoft Foundry, and the MAI Playground
Features of MAI-voice-2.1 models
- ▪Both MAI-Voice-2.1 and MAI-Voice-2.1-Flash models can clone a voice from a few seconds of reference audio, with built-in safeguards to prevent misuse
- ▪MAI-Voice-2.1 supports 23 languages and 26 locales, allowing it to speak multiple languages in the same voice with a native accent
- ▪Microsoft AI released two text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash, alongside the streaming transcription model, MAI-Transcribe-2-Streaming, on October 1, 2026
- ▪MAI-Voice-2.1-Flash is designed for lower latency and cost, hitting a latency of 150 milliseconds and generating 45 seconds of audio
Pricing of the new models
- ▪MAI-Transcribe-2-Streaming's introductory rate of $0.54 per audio hour through the end of 2026 is higher than streaming rates from xAI's Grok Voice Transcribe 2.0 and Meta's Muse Voice Transcribe, compared to $0.10 per hour for the batch MAI-Transcribe-2 model
- ▪MAI-Voice-2.1 is priced at $22 per million characters, while the MAI-Voice-2.1-Flash variant is priced at $15 per million characters
Strategic goals for the MAI family
- ▪Microsoft AI aims to use the MAI model family, including MAI-Transcribe-2-Streaming, to power its Copilot agents in platforms such as Excel and Outlook
- ▪Microsoft AI Chief Executive Mustafa Suleyman instructed researchers to focus on the MAI model family, including MAI-Transcribe-2-Streaming, to reduce and ultimately eliminate costs paid to Anthropic and OpenAI
Debatable claims
- ▪Real-time streaming AI is too resource-intensive to justify its minor latency improvements
- ▪AI transcription models are accurate enough to fully replace human transcribers
- ▪The commercial benefits of voice-cloning technology justify its potential for misuse
Story comments
Loading comments…