ElevenLabs has launched ElevenLabs v4 and v4 Turbo, introducing a new speech model architecture that supports over 90 languages and enables voice cloning with just 10 seconds of audio. The release focuses on enhanced emotional expression control via sequential inline tags and improved voice consistency across long-form content. The Turbo variant achieves a low latency of 150 milliseconds for real-time voice agents. This launch comes as ElevenLabs' annualized recurring revenue reaches $600 million, with over 55% of its business driven by large enterprise clients.
Model release and availability
- ▪ElevenLabs launched two new speech models, ElevenLabs v4 and ElevenLabs v4 Turbo, on September 28, 2026, featuring improved expression control, lower latency, and a new architecture.
- ▪ElevenLabs made ElevenLabs v4 and ElevenLabs v4 Turbo available in ElevenAgents, ElevenCreative, and through its API.
Improved expression and dialogue control
- ▪ElevenLabs v4 expands inline script tags, allowing users to stack multiple tags in sequence to define emotions, pauses, and sound effects like laughter, whispers, and slamming doors.
- ▪ElevenLabs' new v4 speech model allows AI speakers in dialogue to respond to the context of an entire scene rather than delivering each line in isolation
- ▪ElevenLabs' new v4 speech model maintains voice identity and pacing consistency across long productions up to 10,000 characters per request, preventing voice drift during line regenerations and segment transitions
Expanded multilingual capabilities
- ▪ElevenLabs v4 expanded its language support to more than 90 languages, up from 70 languages in ElevenLabs v3, with the largest quality improvements observed in Japanese, Brazilian Portuguese, Mandarin, and Cantonese.
- ▪ElevenLabs v4 allows cloned voices to speak other languages with native accents without drifting back to their original accents over time.
Voice cloning capabilities
- ▪ElevenLabs v4 reintroduces support for Professional Voice Clones, which were unsupported in the ElevenLabs v3 model.
- ▪The ElevenLabs v4 model architecture allows users to clone a voice using only 10 seconds of audio.
Performance benchmarks
- ▪ElevenLabs v4 scored 91.7 percent on the pronunciation benchmark, representing an improvement from the 85.6 percent score of ElevenLabs v3.
- ▪In ElevenLabs' blind tests, listeners rated ElevenLabs v4 as more expressive in 65% to 81% of comparisons against competing models from Cartesia, Inworld, and Google.
ElevenLabs v4 Turbo latency and optimization
- ▪The ElevenLabs v4 Turbo model starts producing audible speech in 150 milliseconds, compared to 262 milliseconds for Cartesia Sonic 3.6 and 814 milliseconds for OpenAI's GPT-4o mini TTS.
- ▪ElevenLabs v4 Turbo is optimized to start generating audio as soon as the underlying Large Language Model begins generating text answers, enabling more fluid conversations for voice agents.
Model pricing
- ▪ElevenLabs introduced temporary promotional API pricing through October 12, 2026, reducing ElevenLabs v4 to $22 per million characters and ElevenLabs v4 Turbo to $11 per million characters.
- ▪The standard API pricing for ElevenLabs v4 is $80 per million characters, while ElevenLabs v4 Turbo costs $40 per million characters.
ElevenLabs business growth and revenue
- ▪More than 55% of ElevenLabs' business revenue comes from large enterprise clients, including Klarna, Deutsche Telekom, Cisco, and Adobe.
- ▪ElevenLabs' annualized recurring revenue reached $600 million in September 2026, up from approximately $330 million at the start of the year.
Debatable claims
- ▪Licensing AI voice clones does more harm than good for professional voice actors
- ▪A ten-second audio sample is insufficient to prevent malicious voice cloning
Story comments
Loading comments…