This page was machine-translated and may differ from the original. View original
Supporting Over 90 Languages, Average 100ms Response, Ranks First in Independent Benchmarks
ElevenLabs has simultaneously launched 'Eleven v4' and 'Eleven v4 Turbo,' text-to-speech (TTS) models based on a new architecture. Both models simultaneously implement expanded emotional expression range and improved response speed, achieving top rankings in the independent benchmarking organization Artificial Analysis' Provider Voice Arena evaluation for September 2026.
ElevenLabs unveiled the two models on the 1st, with co-founder and CEO Mati Staniszewski stating, "In many areas of economy and society, the way AI communicates with people is as important as its reasoning capability. Eleven v4 provides the rich expressive ability that has been lacking in AI, and will elevate artistic creation and agent experiences in more markets worldwide," he added.
Co-founder and Chief Technology Officer Piotr Dabkowski stated, "Based on a completely new architecture, it has broad emotional expression and fine-grained control capabilities, and was designed with low-latency applications in mind, providing the same expressive power in the conversational agent field."
Eleven v4 Strengthens Expressiveness and Consistency
According to the company, Eleven v4 interprets tone, speed, and context contained in text to deliver emotionally vivid speech ranging from urgent atmospheres to warm or witty tones.
Users can specify expression style in natural language or fine-tune utterance style and acoustic effects using inline tags such as '[laughter]' and '[light rain].'
Speaker consistency has improved compared to the previous Eleven v3 generation, and performance in implementing multi-party conversations has also been enhanced.
Supported languages have increased from 70 to over 90, with notably improved performance in Japanese, Mandarin Chinese, and Brazilian Portuguese.
It also provides functionality to switch language and intonation while maintaining the pitch and timbre of the original voice.
Eleven v4 Turbo Targeting Real-Time Agents
According to the company, Eleven v4 Turbo has an average intermediate inference latency of 100 milliseconds (ms) until the first response is generated, delivering the same emotional expressiveness of Eleven v4 even in real-time conversational environments.
It can be applied to conversational agents in diverse industries including healthcare, finance, and gaming, and is optimized as a single system with ElevenLabs' conversational agent platform 'ElevenAgents.'
In terms of security, it complies with global regulations including the EU General Data Protection Regulation (GDPR), the US Health Insurance Portability and Accountability Act (HIPAA), SOC 2, and PCI, with voice cloning utilizing 'Voice CAPTCHA' authentication and 'No-Go Voices' safeguards that block major public voice replications.
Both models are available through ElevenCreative, ElevenAgents, and ElevenAPI.
Meanwhile, cumulative revenue generated by creators who have registered voices in ElevenLabs' 'Voice Library' has exceeded $22 million (approximately 30 billion Korean won) to date.
To request a correction, reply or follow-up report on this article, see how to file a request. Previously published statements are collected in corrections & replies.


















