ElevenLabs launched Eleven v4 and v4 Turbo speech models, offering more accurate direction cues for laughter and whispers, consistent voice quality across long productions, and support for over 90 languages. The Turbo variant delivers 150-millisecond latency for real-time voice agents.
Google DeepMind launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, 2026, its most advanced native speech-to-speech models for real-time voice agents. At $1.38 per hour of voice conversation, the models significantly undercut OpenAI's GPT-Live-1 pricing while topping the Artificial Analysis speech-to-speech leaderboard.
Jason Isbell and three other artists filed a lawsuit against AI music company Suno, accusing it of training its model on their songs and imitating their voices. The suit tests a legal strategy based on unlawful imitation rather than copyright.
Meta Superintelligence Labs has released its first product, Muse Voice Transcribe, a real-time AI model that combines speech transcription, speaker diarization, and endpointing in a single system. The model can distinguish between multiple speakers and languages simultaneously, replacing the traditional approach of using three separate systems.
Plaud Inc. has unveiled the Plaud One Explorer, wireless earbuds that can capture and transcribe conversations, with a charging case featuring an embedded eSIM that enables AI-assisted tasks like drafting documents without requiring a phone connection.
Google has launched Gemini 3.5 Transcribe, a new AI-powered speech-to-text model that automatically removes filler words and cleans up speech in real-time. The model reports a 2.6% average word error rate across more than 85 languages and will be integrated into Chrome's voice typing feature.
Ringg, an Indian startup developing voice AI agents for business communications, has secured $10 million in funding from Peak XV Partners. The company automates routine business conversations across industries, from clinic bookings to abandoned cart follow-ups, capitalizing on India's strong consumer preference for phone-based communication.
Google's Gemini artificial intelligence assistant has crossed 1 billion monthly active users, making it the fastest-growing product in the company's history. The milestone comes as 63% of users engage with the voice feature and the platform generates over 150 million images daily.
Anthropic announced an update to Claude's voice mode that adds support for the more capable Opus and Sonnet models alongside Haiku. The update also allows users to switch models mid-conversation and use connected apps and multiple languages.
OpenAI unveiled Presence, a new enterprise platform for deploying AI voice agents and chatbots, while announcing plans to spend $750 billion on infrastructure through 2030—a 25% increase from earlier estimates. The moves come as the company faces pressure to prove AI's business value, while Emarketer projects that OpenAI will fall far short of its AI advertising revenue targets.
OpenAI launched GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper in its Realtime API, enabling developers to build reasoning agents, speech translation across 70+ languages, and streaming transcription capabilities.
Folk artist Murphy Campbell discovered multiple AI-generated fake songs attributed to her on Spotify, becoming a target for both AI impersonation and copyright trolling in a case highlighting vulnerabilities in music platforms' content verification systems.
Microsoft's MAI division released three new foundation models: MAI-Transcribe-1 (speech-to-text in 25 languages, 2.5x faster than Azure Fast), MAI-Voice-1 (audio generation), and an image generation model. The transcription model costs $0.36 per audio hour.