ElevenLabs' New Voice Models Just Took the Top Two Spots on a Major Leaderboard
Kaino
9h agoOct 6, 2026, 12:00 AM5 views

ElevenLabs' New Voice Models Just Took the Top Two Spots on a Major Leaderboard

ElevenLabs has released Eleven v4 and v4 Turbo — one built for expressive speech, one for low-latency voice agents. Both top the Artificial Analysis leaderboard.

ElevenLabstext-to-speechvoice AIEleven v4voice agent latencyArtificial Analysis leaderboardAI speech models

ElevenLabs has released two new text-to-speech models built for deliberately different jobs. Eleven v4 is positioned as the company's most expressive flagship model. Eleven v4 Turbo is built for speed, targeting fast conversational responses in voice agents. That split itself is the real news here — it's an explicit bet that expressive narration and responsive conversation need different models, not one model trying to do both.

The tension it's addressing is real and familiar in voice AI: a narrated character or accessibility reader benefits from deliberate pacing and emotional delivery, while a voice agent needs to start talking fast enough to not feel broken. TechCrunch confirms both models add expanded expression controls and support more than 90 languages, with Turbo specifically positioned as the low-latency option for voice agents.

Turbo's speed claim is genuinely concrete, which is worth noting since most AI speed claims aren't. ElevenLabs reports roughly 100ms median inference latency and about 150ms median time to first audible speech. The second number is the one that actually matters to a person using the product — it's the gap between asking something and hearing a reply start, not just how fast the model processes internally. What's not established: how those medians were measured across different prompt lengths, languages, or concurrent traffic, and what the full end-to-end latency looks like once you add speech recognition, application logic, and networking on top of TTS generation alone.

Both models immediately landed the top two spots on Artificial Analysis' Provider Voice leaderboard — Turbo first at 1334±19 Elo, v4 second at 1321±18. That's a genuinely strong public result under one specific methodology. It's not proof either model is the right choice for every use case, though — the leaderboard doesn't speak to long-form narration quality, specific accents, voice cloning, cost, or how performance holds across all 90+ supported languages individually.

Bottom line: ElevenLabs has made a clear, coherent two-model bet — quality versus speed, with both options scoring well on an independent leaderboard right out of the gate. What's still unproven is whether Turbo's impressive published latency survives inside a real production voice-agent stack, and whether expression and language quality stay consistent across every language ElevenLabs now claims to support.

Key takeaways
  • 1

    ElevenLabs has released two new text to speech models built for deliberately different jobs.

  • 2

    Eleven v4 is positioned as the company's most expressive flagship model.

  • 3

    Eleven v4 Turbo is built for speed, targeting fast conversational responses in voice agents.

Continue reading

Latest from Kaino News