Klatt

BlogResearch

On Voice AI Rankings

What a Spanish listening test tells us about voice models

Which voice model is the best?

Voice preferences, like preferences for paintings or wine, have a difficult to quantify element. Mean Opinion Scores and Elo-style rankings turn those judgments into numbers, but understanding what those numbers mean or their relevance for a particular case takes a closer look.[1]

There is good work on this outside academia. Artificial Analysis, for example, evaluates both controlled cloned voices and providers’ own voices. We were interested in a narrower question however, using sentences relevant to our work, and being able to fully understand the rankings.[2]

So we built the comparison we wanted to use. You can try the live listening survey and review the code, and analysis on GitHub.

Two recordings and one choice

The survey compared seven systems: Qwen3-TTS, VoxCPM2, Chatterbox, IndexTTS, Fish, F5-Spanish and ElevenLabs v3. This was a practical selection of open-weight models with cloning capabilities, with ElevenLabs as our commercial reference point.[3]

Each received the same customer-service texts to work on and the same clip from the reference speaker to clone.

We skipped standard voices to reduce differences caused by voice selection, and limited ourselves to asking survey participants to decide between ten sets of two voice pairs with only one question: ¿Which sounds more natural?

Competitive results and an awkward ranking

On 17 September 2026, the survey had 326 judgments. Qwen led the overall ranking, with ElevenLabs close behind. The ranking combines results against every opponent, rather than simply counting who won their direct encounter.

Figure 1 A narrow lead with substantial uncertainty

Voice model benchmark

Text-to-speech leaderboard

7 systems / Ranked by rating

Text-to-speech leaderboard
RankSystemRatingWins / trials
01
Qwen3-TTSQwen · Alibaba Cloud
1582.959 / 93
02
ElevenLabs v3ElevenLabs
1563.956 / 93
03
VoxCPM2OpenBMB
1545.054 / 94
04
F5-Spanishjpgallegoar · F5-TTS
1482.744 / 93
05
ChatterboxResemble AI
1464.841 / 93
06
Fish S2 ProFish Audio
1457.140 / 93
07
IndexTTS 2.5IndexTTS · Bilibili
1403.632 / 93

Higher rating ranks first · One speaker, 15 Spanish prompts

About 65%: Model-estimated probability that Qwen’s underlying strength exceeds ElevenLabs’

The interesting part appears when we look at individual matchups and how each pair compared when set up against the other across multiple parings.

Figure 2 : Face to Face preferences amongst providers
ElevenLabsQwen

ElevenLabs 8, Qwen 8.

QwenF5-Spanish

Qwen 9, F5-Spanish 6.

F5-SpanishElevenLabs

F5-Spanish 9, ElevenLabs 6.

Blue = Win, Grey = Lose

Qwen and ElevenLabs were tied at 8–8. Qwen beat F5 by 9–6, while F5 beat ElevenLabs by the same margin, unveiling a fuzzy reality hidden behind the single number.

Is this fair?

Every system received the same 11.74-second reference recording, significantly less that recommended by some providers.[4]

We used a public-domain recording of a male audiobook narrator as our reference voice. Audiobook narration is not an ideal proxy for natural customer-service speech, and testing a single cloned voice leaves out much of each system’s range. Given the legal complexities around voice use, this was the most practical reference we could find.

Given the burden is shared equally however, we considered it would still be informative.

The question itself also matters. We asked which recording sounded more natural, not which was best or which someone would choose for their business. The experiment assumption that “more natural = better” is in itself debatable.

Curiously also, presentation order appeared to matter too. Across all 326 judgments, the assigned-second option (Option B) won 61.0% of the time. Given the balancing of order appearance and multiple rounds among providers however, we have done our best to neuter this effect and the figures are fully transparent.

The main issue: A voice model alone is not a product

While the open-weight models were competitive in our test, the lesson for those in the voice AI market is less straightforward. Voice itself is only part of what customers are buying.

Businesses buying voice AI need a complete system that solves a business need. That means bringing an adequate voice, speech recognition, a language model and speech generation together with low enough latency for a natural conversation, and reliable turn-taking and barge-in so callers can interrupt. It also means tool calling and direct integration with existing systems, so the agent can access relevant information and take action.[5][6]

This without beginning to consider aspects like infrastructure reliability, monitoring, human handover, security, compliance with local legislation and data sovereignty.

A complete voice product

Models, integrations and operations working towards a business outcome

CONVERSATION CONTROL

End-to-end latencyTurn-takingBarge-in

Response time across models, network and tools; callers can interrupt

MODELS AND VOICE

Speech recognition

Speech → text

Language and accent coverage

Language model

Decides what to say or do

Context, instructions and reasoning

Speech generation

Text → speech

Voice choice and pronunciation

Tool calling and direct integration with existing systems

CRM • Orders • Bookings • Knowledge bases • Business APIs

Retrieve information, take action and return results to the conversation

RUNNING THE SERVICE

Reliable operations
  • Infrastructure, capacity and availability
  • Testing, monitoring and error recovery
  • Human handover with conversation context
Security and control
  • Security, privacy and access controls
  • Compliance with local legislation
  • Data sovereignty: hosting and data control
Commercial fit

Usage, infrastructure and maintenance costs • Ability to change providers

A common modular architecture; some systems combine model stages.

Our listening test evaluated generated audio, not this complete product.

Matching a commercial provider’s voice in a listening test does not mean matching its product. A business may be willing to pay more for an integrated service because it saves development work, is easier to operate or helps resolve more customer requests successfully.

But how much more? If several voice models sound good enough for the bounded task, it will become increasingly harder to justify a premium on voice quality alone. For example, Fish Audio lists S2.1 Pro at USD 15 per million UTF-8 bytes, while ElevenLabs lists Flash/Turbo at USD 50 per million characters. Those units are not interchangeable, so the actual cost difference depends on the text.[7][8]

Perhaps it is wise to not be tied to a single provider for the future, but one which will be specific to solve your business needs, and to a platform that will give you a choice.

If that is the case, at Klatt we would love to hear from you.

Data and methods

Results use a saved public snapshot generated on 17 September 2026: 326 judgments from 36 voting browser sessions, 31 complete. Sessions are not verified unique people. Seven systems, 15 fixed sentences, one reference speaker and one first completed take per model/text. Results remain provisional.

Ratings use a regularized Bradley–Terry model on an Elo-style scale. The refreshed aggregate Bayesian fit estimates a 65.4% probability that Qwen’s underlying strength exceeds ElevenLabs’. It uses Normal(0,2) strength priors, four chains and 12,000 retained draws, with no divergences and maximum parameter R-hat 1.007 after rounding up. It cannot adjust for individual sessions, text or order. Qwen’s fitted lead is 19.0 rating points; removing F5 changes it to −2.0, while removing IndexTTS leaves +4.4. These results replace the earlier 319-judgment aggregate analysis.

Position and duration analyses use a read-only snapshot of all 326 judgments, captured at 17:59:25 Madrid time on 17 September 2026. The assigned-second option won 199/326 comparisons (61.0%). Updated models accounting for model, clip and session differences still found a position association after including duration; evidence for a preference for longer clips remained inconclusive. Duration is observational and confounded with pacing, pauses and model identity. Because 325/326 trials presented A first, position cannot be separated from the B label. Trial names were hidden, but named aggregate results were available during participation. Human audio QA remains pending. This pilot does not establish listener representativeness for Spain or correctness of every utterance.

Pricing checked 18 September 2026 against official provider pages. The examples cover speech synthesis only. Fish bills UTF-8 bytes, which are not interchangeable with characters. The pricing examples include models outside this survey. We did not measure self-hosting costs.[7][8]

Legal detail: consent is one possible GDPR lawful basis, not the only one. Voice is not automatically special-category biometric data; that restriction concerns biometric processing for unique identification. AI Act Article 50 distinguishes provider marking duties from deployer disclosure of qualifying deepfakes, with exceptions and applicable transition rules.[9][10]

The GitHub repository contains the frozen results and instructions for reproducing the numerical analyses; recordings are excluded. The live survey may show newer totals than this article’s dated snapshot.

Sources and further reading

  1. Kayyar Lakshminarayana et al. (2023): absolute ratings and ranking-based TTS evaluation. Useful rationale; no universal superiority of relative tests. ↩
  2. Artificial Analysis: TTS benchmark methodology. Controlled voices, provider voices and language-specific comparisons. ↩
  3. Qwen3-TTS official repository. Open-weight models and reference-audio voice cloning. ↩
  4. ElevenLabs: Instant Voice Cloning. Reference-audio recommendations. ↩
  5. ElevenAgents documentation. Architecture, integrations and operating tools. ↩
  6. Pipecat documentation. Open-source orchestration and deployment options. ↩
  7. ElevenLabs API pricing. Published character rates; checked 18 September 2026. ↩
  8. Fish Audio developer pricing. S2.1 Pro paid API rate per UTF-8 byte; checked 18 September 2026. ↩
  9. GDPR, Articles 4–6 and 9. Personal data, lawful processing and biometric data (official Spanish text). ↩
  10. EU AI Act, Article 50. Transparency obligations for providers and deployers. ↩