Research
Spanish customer-service TTS benchmark
A dated, limited comparison of seven systems using 326 listening judgments.
Snapshot facts
- Judgments
- 326
- Systems
- 7
- Prompts per system
- 15
- Reference speaker
- One shared 11.74-second reference speaker
- Snapshot date
- 2026-09-17
| Rank | System and version | Rating | Wins / trials |
|---|---|---|---|
| 1 | Qwen3-TTS 1.7BQwen/Qwen3-TTS-12Hz-1.7B-Base@fd4b254389122332181a7c3db7f27e918eec64e3 | 1582.9 | 59 / 93 |
| 2 | ElevenLabs v3eleven_v3 | 1563.9 | 56 / 93 |
| 3 | VoxCPM2openbmb/VoxCPM2@32279effe8c19989596f05d353d1447f51d9e915 | 1545.0 | 54 / 94 |
| 4 | F5-Spanishjpgallegoar/F5-Spanish@4765c14ffd01075479c2fde8615831acc0adca9a | 1482.7 | 44 / 93 |
| 5 | Chatterbox SpanishResembleAI/Chatterbox-Multilingual-es-es@7262b364e3222ba9446c9289e7485eadc47ed902 | 1464.8 | 41 / 93 |
| 6 | Fish Audio S2 Profishaudio/s2-pro@1de9996b6be38b745688de084d87a5633f714e4e | 1457.1 | 40 / 93 |
| 7 | IndexTTS 2.5IndexTeam/IndexTTS-2.5@c39ce5ba981572cb187443877ff559dfb246ce63 | 1403.6 | 32 / 93 |
The 65.4% fitted-strength probability describes uncertainty in this aggregate ordering. It is not a quality percentage or a universal best-model claim.
How to read the table
- Rank
- Rank is the position in this dated aggregate ordering.
- Rating
- Rating is the Elo-style presentation of the regularized Bradley–Terry aggregate.
- Wins / trials
- Wins / trials are the observed comparison outcomes in the saved snapshot.
Method
Each system used the same 11.74-second reference speaker, the same 15 Spanish customer-service prompts, the first completed take, and loudness matching. Listeners compared two anonymized recordings with replay available and no tie option.
Regularized Bradley–Terry model on an Elo-style scale. The aggregate uses a Normal(0, 2) strength prior, four chains, and 12,000 retained draws. Maximum parameter R-hat was 1.007 after rounding.
Sensitivity checks: removing F5 changes Qwen’s fitted lead from 19.0 rating points to −2.0; removing IndexTTS leaves a +4.4-point lead. Position and duration remain confounded because 325 of 326 trials presented option A first.
Ratings summarise the tournament, not a percentage score. Qwen and ElevenLabs tied at 8–8 directly, so the aggregate ordering should be read as a qualified lead.
Sources and model versions
- Qwen3-TTS 1.7B https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-BaseOfficial source checked 17 September 2026
- VoxCPM2 https://huggingface.co/openbmb/VoxCPM2Official source checked 17 September 2026
- Chatterbox Spanish https://huggingface.co/ResembleAI/chatterboxOfficial source checked 17 September 2026
- IndexTTS 2.5 https://github.com/index-tts/index-ttsOfficial source checked 17 September 2026
- Fish Audio S2 Pro https://fish.audioOfficial source checked 17 September 2026
- F5-Spanish https://huggingface.co/jpgallegoar/F5-SpanishOfficial source checked 17 September 2026
- ElevenLabs v3 https://elevenlabs.io/text-to-speechOfficial source checked 17 September 2026
Limitations
One shared speaker, fixed prompts, and a first-take protocol narrow the question. Human audio QA is pending, listeners are not verified representative of Spain, and the survey did not measure live streaming or full production performance.