Mon, 28 Sep

ElevenLabs Releases Eleven v4 and v4 Turbo Neural Networks: AI Now Follows Directorial Commands and Clones a Voice in 10 Seconds

Max Ivanov · 28.09.2026 18:12 · 3 min read

ElevenLabs has unveiled a new generation of speech synthesis models — Eleven v4 and its high-speed variant, Eleven v4 Turbo. The updated architecture fundamentally changes the approach to machine voiceover. Where algorithms once simply read text while trying to guess the right intonation on their own, now the user takes on the role of director. For audiobook creators, video game developers and content creators, this means less time spent on endless regenerations to get the right take.

The headline feature is a system for controlling emotions through text prompts. On the official release page, the developers explain that stage directions in square brackets can be inserted right into the script: [whispers], [shouting], [laughing], [sighs] or [long pause]. The model doesn’t just produce the specified sound — it organically adjusts the entire pace and manner of speech to the given emotion. The algorithm has also learned to generate synchronized sound effects alongside speech — for example, the sound of a door slamming in the background.

Context stitching for long texts and 90 languages

The base Eleven v4 supports more than 90 languages, including Russian and Ukrainian, and handles multi-voice dialogues well.

Specifically for publishers and audiobook authors, the workspace now includes context stitching technology. The algorithm solves a long-standing problem of neural voiceover for long formats: it links individual generated fragments together, preserving a consistent intonation, rhythm and character of the narrator’s voice throughout an hours-long recording.

Working with custom voices has also become easier. To create a quick digital copy (Instant Voice Cloning), the system now needs just 10 seconds of clean source audio. In addition, as stated in the platform’s documentation, the fourth generation brings back full support for studio-quality Professional Voice Cloning, which was missing in the third version.

The Turbo version and natural ping for AI agents

For businesses and developers of interactive services, ElevenLabs released Eleven v4 Turbo. This model keeps the functionality of the base version but is extremely optimized for real-time operation. Median inference latency (the neural network’s response time) has been reduced to 100 ms, and no more than 150 ms passes before the first sound is spoken. This is a critical metric for voice assistants, support services and phone bots — such latency isn’t perceived by the human ear as a pause, making dialogue with the machine as natural as possible.

Both generative models are already deployed in the company’s infrastructure and are available through the ElevenCreative and ElevenAgents interfaces and via API. You can test Eleven v4’s functionality on the basic free tier, which includes 10,000 generation credits per month.

Enjoy VseZavislo?

Add us to your preferred Google sources to see our news more often.

Add us to your Google

Share

Leave a comment