ElevenLabs launches v4 with faster voice cloning and expression control

ElevenLabs launches v4 with faster voice cloning and expression control

ElevenLabs has launched v4 and v4 Turbo, two new speech models focused on expressive generation and lower latency. v4 uses a new architecture designed to improve vocal control, voice consistency and cloning speed, with ElevenLabs saying it can clone a voice from about 10 seconds of audio. It also considers broader text context when adjusting delivery and expression across longer passages.

Building on the inline expression tags introduced with v3, v4 lets users stack multiple tags and follows them sequentially for more granular control over delivery. The new generation also expands language support from around 70 to more than 90 languages, with ElevenLabs reporting some of the largest quality improvements in Japanese, Brazilian Portuguese, Mandarin and Cantonese.

v4 Turbo is aimed at latency sensitive use cases such as conversational voice agents, prioritizing faster responses, while the standard v4 model focuses more on expressive speech and longer form content.

by Mauricio B. Holguin

Add as a preferred source on Google
  • ...

Leverages advanced AI for natural speech, offering applications like podcasts, video voiceovers, with a user-friendly interface and extensive voice library.

No comments so far, maybe you want to be first?