ElevenLabs launches v4 and Turbo, topping TTS benchmark ahead of Gemini and Qwen

ElevenLabs launches v4 and Turbo, topping TTS benchmark ahead of Gemini and Qwen

N
News Editor
2026-09-29 10:37:02
ElevenLabs has released its new voice models, Eleven v4 and the real-time focused Eleven v4 Turbo, with the flagship model moving to the top of the latest Artificial Analysis text-to-speech blind test ranking. Eleven v4 is listed at 1315 Elo, ahead of Cartesia Sonic 3.6 at 1275, Gemini 3.8 Flash TTS at 1267, and Qwen-Audio-3.0-TTS-Plus at 1258. The company’s previous-generation Eleven v3 currently stands at 1169 Elo, placing 17th in the same ranking. The new release centers on tighter control over tone, pacing, emotion, and sound effects. According to the source material, users can either specify those parameters directly or describe in natural language how a sentence should be spoken. ElevenLabs also expanded language coverage from more than 70 languages to 90-plus languages, while Instant Voice Clone requires about 10 seconds of audio. The company said the update also improves voice consistency in long-form generation and multi-speaker dialogue. For latency, official documentation shows median inference time for v3 Conversational at about 280ms, while v4 Turbo cuts that to roughly 100ms for real-time voice agent use cases.

ElevenLabs has introduced its next-generation voice model, Eleven v4, alongside Eleven v4 Turbo, a version built for real-time voice agents.

Eleven v4 leads the latest TTS blind ranking

In the latest text-to-speech blind test ranking from Artificial Analysis, Eleven v4 is currently ranked No. 1 with a score of 1315 Elo. Cartesia Sonic 3.6 is listed at 1275, Gemini 3.8 Flash TTS at 1267, and Qwen-Audio-3.0-TTS-Plus at 1258.

The previous-generation Eleven v3 is currently at 1169 Elo, placing 17th.

v4 focuses on more precise control over delivery

Eleven v4 is designed to deliver more natural emotional expression and finer voice control. The source says v3 already supported audio tags such as laughter and whispering, while v4 mainly improves the precision of that control system.

Users can directly specify tone, pacing, emotion, and sound effects, or use natural language to describe how a line should be spoken. The model also expands language coverage from more than 70 languages to more than 90.

ElevenLabs said Instant Voice Clone requires about 10 seconds of audio. The update also strengthens voice consistency in long-form content and multi-speaker conversations.

Turbo cuts latency for real-time conversations

Eleven v4 Turbo is aimed at speed in live conversational settings. According to official documentation cited in the source, median inference latency for v3 Conversational is about 280ms, while v4 Turbo reduces that figure to about 100ms.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
1800

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.