Table of Contents
ElevenLabs has launched Eleven v4, its latest text-to-speech model, alongside Eleven v4 Turbo, a lower-latency version designed for real-time voice applications.
The new models expand ElevenLabs’ speech technology with support for more than 90 languages, greater control over vocal expression and voice-cloning capabilities. The company says the models are designed for applications ranging from creative content and narration to conversational AI and voice agents.
Eleven v4 Supports More Than 90 Languages
One of the major updates in Eleven v4 is expanded language support.
ElevenLabs says the model supports more than 90 languages, with improvements highlighted in languages including Japanese, Brazilian Portuguese, Mandarin and Cantonese.
The broader language coverage could support applications such as:
- Audiobooks and narration
- Video and voice content
- Dubbing
- Gaming
- Education
- AI assistants
- Conversational applications
Actual performance can vary depending on the language, voice and application.
More Control Over AI Voice Expression
ElevenLabs is positioning Eleven v4 around more expressive and controllable speech generation.
The model can interpret contextual instructions within scripts, allowing creators to guide how lines are delivered. Audio tags can also be used to indicate directions such as laughter or whispering.
These controls are intended to give creators more flexibility than conventional text-to-speech systems.
ElevenLabs describes Eleven v4 as its most expressive speech model so far. This is the company’s own characterization rather than an independent industry ranking.
Voice Cloning With Short Audio Samples
Another feature highlighted with Eleven v4 is voice cloning.
ElevenLabs says its technology can create a voice clone using as little as 10 seconds of audio. The company also says its models are designed to maintain speaker characteristics during longer generations.
This could be useful for narration, character voices and other projects that require consistent synthetic speech.
Voice cloning also requires appropriate consent. Users should have permission before creating or deploying a replica of another person’s voice.
Eleven v4 Turbo Targets Real-Time AI
Alongside Eleven v4, ElevenLabs has launched Eleven v4 Turbo.
The Turbo model is designed for applications where speech needs to be generated quickly, particularly during live conversations. It supports streaming generation, allowing audio to begin playing before an entire response has been processed.
ElevenLabs reports a median time to first audio of around 150 milliseconds in its testing, excluding network latency. This figure should not be treated as a guaranteed response time for every application.
The company also lists approximately 100 milliseconds of median inference latency for the Turbo model under its documented testing conditions.
Potential applications include:
- AI voice agents
- Virtual assistants
- Customer-service systems
- Interactive characters
- Real-time conversational tools
Eleven v4 Enters a Competitive AI Voice Market
The launch comes as AI voice technology continues to develop rapidly.
Companies including OpenAI, Google and specialist voice-AI providers are developing speech-generation and conversational systems. Developers evaluating these technologies typically consider factors such as voice quality, language coverage, expression control, latency and integration options.
Eleven v4 focuses on expressive speech and control, while Eleven v4 Turbo places greater emphasis on speed for real-time applications.
The best option can therefore depend on the requirements of a particular project rather than one single performance measurement.
What Eleven v4 Means for Developers
For developers, Eleven v4 provides another option for building applications that depend on AI-generated speech.
The standard model is designed for expressive speech generation, while Eleven v4 Turbo is intended for applications where faster responses are important.
The two models can therefore address different requirements, from long-form narration and creative content to interactive voice agents.
FAQs
What is Eleven v4?
Eleven v4 is ElevenLabs’ latest text-to-speech model. It supports more than 90 languages and is designed to provide greater control over expressive AI speech.
What is Eleven v4 Turbo?
Eleven v4 Turbo is a lower-latency model designed for real-time speech applications such as AI voice agents and conversational assistants.
How many languages does Eleven v4 support?
ElevenLabs says Eleven v4 supports more than 90 languages, including Japanese, Brazilian Portuguese, Mandarin and Cantonese.
Can Eleven v4 clone voices?
Yes. ElevenLabs says its voice-cloning technology can create a clone using as little as 10 seconds of audio. Appropriate permission should be obtained before cloning another person’s voice.
What can Eleven v4 be used for?
Eleven v4 is designed for applications including narration, audiobooks, dubbing, gaming, content creation and conversational AI.
Conclusion
ElevenLabs’ Eleven v4 expands its text-to-speech technology with support for more than 90 languages, voice-cloning capabilities and additional controls for expressive speech.
The accompanying Eleven v4 Turbo focuses on lower-latency generation for real-time applications such as AI voice agents and conversational assistants.
As AI voice technology becomes increasingly competitive, developers will need to consider language support, speech quality, latency and integration requirements when selecting a model. Eleven v4 gives creators and developers another option for building both creative and interactive voice experiences.

