we_are_coded.by CODE · The world, decoded
БГ
ElevenLabs

ElevenLabs releases Eleven v4 - a voice that reads the tone of a sentence, not just the words

ElevenLabs · event date: 28 September 2026Audio

On 28 September ElevenLabs introduced Eleven v4, an entirely new text-to-speech architecture, and the fast Eleven v4 Turbo for voice agents. Per the company the model ranks first on the Artificial Analysis leaderboard, and in blind comparisons about 75 percent of listeners prefer it.

In short
  • Delivery is directed with natural-language descriptions and tags such as [laughs] or [light rain].
  • Turbo is built for agents: median inference latency around 100 milliseconds, per the company.
  • An instant voice clone needs 10 seconds of audio. Both models support more than 90 languages.
Checked on1 October 2026Responsible editorTsvetelin IvanovHow we workMethod · Corrections

"I need you to stay calm." A doctor says it quietly to a frightened patient. A character in a game shouts it at his squad. The words are the same.

ElevenLabs opens its announcement with that example, and it is right to. A voice that reads the words correctly stopped being news a while ago; the news is a voice that understands how a sentence should be said.

Anyone who has stood behind the desk at an event knows that the same "good evening" can lift a room or put it to sleep. The difference is never in the text.

The facts: on 28 September 2026 ElevenLabs released Eleven v4 and the low-latency Eleven v4 Turbo. Per the company v4 is built on an entirely new architecture, is ranked first by Artificial Analysis (Provider Voice Arena, September 2026) and was preferred by about 75 percent of listeners in blind comparisons against Cartesia Sonic 3.6, Inworld TTS-2 and two Google Gemini 3.8 voice models. Turbo has a median inference latency of about 100 milliseconds and a median time to first speech of about 150 milliseconds, and was built together with the ElevenAgents agent platform. Delivery can be directed with natural-language descriptions and inline tags such as [laughs] or [light rain]. Both models support more than 90 languages, and an instant voice clone needs 10 seconds of audio. They are available in ElevenAgents, ElevenCreative and through ElevenAPI.

The upside is obvious: an audiobook with real characters, dubbing where a brand sounds like itself in every language, a voice agent that does not sound like a machine having a bad day.

Ten seconds of your voice is enough.

The other side is just as obvious. Ten seconds of audio for a clone is great news for dubbing and unwelcome news for anyone whose voice is somewhere online. And if a voice can perform a tone, it can perform someone else's tone too.

I will listen to it in Bulgarian before I say more. More than ninety languages in an announcement is a promise, and whether our stress falls in the right place takes a minute to hear.

The visual is generated code art. No third-party images.
Follow usFacebookLinkedIn
Official primary sources
→ElevenLabs - Introducing Eleven v4, our most emotive model, 28.09.2026
Original: https://wearecoded.com/en/articles/elevenlabs-eleven-v4-tonat.html
ShareFacebookXLinkedInTelegramWhatsApp
← Back to all news