On 28 September ElevenLabs introduced Eleven v4, an entirely new text-to-speech architecture, and the fast Eleven v4 Turbo for voice agents. Per the company the model ranks first on the Artificial Analysis leaderboard, and in blind comparisons about 75 percent of listeners prefer it.
- Delivery is directed with natural-language descriptions and tags such as [laughs] or [light rain].
- Turbo is built for agents: median inference latency around 100 milliseconds, per the company.
- An instant voice clone needs 10 seconds of audio. Both models support more than 90 languages.
"I need you to stay calm." A doctor says it quietly to a frightened patient. A character in a game shouts it at his squad. The words are the same.
ElevenLabs opens its announcement with that example, and it is right to. A voice that reads the words correctly stopped being news a while ago; the news is a voice that understands how a sentence should be said.
Anyone who has stood behind the desk at an event knows that the same "good evening" can lift a room or put it to sleep. The difference is never in the text.
The upside is obvious: an audiobook with real characters, dubbing where a brand sounds like itself in every language, a voice agent that does not sound like a machine having a bad day.
The other side is just as obvious. Ten seconds of audio for a clone is great news for dubbing and unwelcome news for anyone whose voice is somewhere online. And if a voice can perform a tone, it can perform someone else's tone too.
I will listen to it in Bulgarian before I say more. More than ninety languages in an announcement is a promise, and whether our stress falls in the right place takes a minute to hear.