Klatt

Research

Spanish customer-service TTS benchmark

A dated, limited comparison of seven systems using 326 listening judgments.

Snapshot: 17 September 2026 · 326 judgments · one shared speaker · 15 prompts per system

Snapshot facts

Judgments
326
Systems
7
Prompts per system
15
Reference speaker
One shared 11.74-second reference speaker
Snapshot date
2026-09-17
Spanish customer-service TTS benchmark
RankSystem and versionRatingWins / trials
1Qwen3-TTS 1.7BQwen/Qwen3-TTS-12Hz-1.7B-Base@fd4b254389122332181a7c3db7f27e918eec64e31582.959 / 93
2ElevenLabs v3eleven_v31563.956 / 93
3VoxCPM2openbmb/VoxCPM2@32279effe8c19989596f05d353d1447f51d9e9151545.054 / 94
4F5-Spanishjpgallegoar/F5-Spanish@4765c14ffd01075479c2fde8615831acc0adca9a1482.744 / 93
5Chatterbox SpanishResembleAI/Chatterbox-Multilingual-es-es@7262b364e3222ba9446c9289e7485eadc47ed9021464.841 / 93
6Fish Audio S2 Profishaudio/s2-pro@1de9996b6be38b745688de084d87a5633f714e4e1457.140 / 93
7IndexTTS 2.5IndexTeam/IndexTTS-2.5@c39ce5ba981572cb187443877ff559dfb246ce631403.632 / 93

The 65.4% fitted-strength probability describes uncertainty in this aggregate ordering. It is not a quality percentage or a universal best-model claim.

How to read the table

Rank
Rank is the position in this dated aggregate ordering.
Rating
Rating is the Elo-style presentation of the regularized Bradley–Terry aggregate.
Wins / trials
Wins / trials are the observed comparison outcomes in the saved snapshot.

Method

Each system used the same 11.74-second reference speaker, the same 15 Spanish customer-service prompts, the first completed take, and loudness matching. Listeners compared two anonymized recordings with replay available and no tie option.

Regularized Bradley–Terry model on an Elo-style scale. The aggregate uses a Normal(0, 2) strength prior, four chains, and 12,000 retained draws. Maximum parameter R-hat was 1.007 after rounding.

Sensitivity checks: removing F5 changes Qwen’s fitted lead from 19.0 rating points to −2.0; removing IndexTTS leaves a +4.4-point lead. Position and duration remain confounded because 325 of 326 trials presented option A first.

Ratings summarise the tournament, not a percentage score. Qwen and ElevenLabs tied at 8–8 directly, so the aggregate ordering should be read as a qualified lead.

Sources and model versions

Limitations

One shared speaker, fixed prompts, and a first-take protocol narrow the question. Human audio QA is pending, listeners are not verified representative of Spain, and the survey did not measure live streaming or full production performance.