Open TTS Leaderboard: Scalable Evaluation For Multilingual Text-to-Speech And Voice Cloning
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get tech for your team delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Hugging Face has introduced the Open TTS Leaderboard to compare open text-to-speech models using measures of intelligibility, speed and speaker similarity. It adds multilingual and voice-cloning comparisons, plus a listening area for community feedback, but its metrics are not a replacement for human preference tests.

Hugging Face has launched the Open TTS Leaderboard, a new tool for comparing open text-to-speech models on intelligibility, inference speed and speaker similarity. It is designed to make evaluation across languages and voice cloning faster than vote-based arenas, while leaving human listening as an important measure of quality.

The leaderboard uses automated measures to assess different parts of model performance. For intelligibility, it compares the supplied text with a transcript of generated audio using word error rate (WER) and character error rate (CER); the report says transcripts are produced with Qwen3 ASR. Speed is measured using inverse real-time factor for batched offline inference and time-to-first-audio (TTFA) for streaming or batch-size-one tests on an H200 GPU and CPU. A speaker-similarity score compares WavLM embeddings from generated speech and a reference clip.

The default ranking uses a macro-average of English WER across the English splits of Seed TTS Eval and CV3 Eval in zero-shot settings. The report names hexgrad/Kokoro-82M, Supertone/supertonic-3 and fishaudio/s2-pro among the leading models on that English measure. Users can switch languages, and a voice-cloning view adds speaker similarity comparisons for models supporting the selected language. The report identifies k2-fsa/OmniVoice, fishaudio/s2-pro and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 as strong multilingual models.

There is also a “Listen” tab where users can compare generated samples, select languages and datasets, and provide feedback. The leaderboard’s authors say community votes may be incorporated as they accumulate. According to the report, automated evaluation can reduce an assessment from weeks of collecting votes to a few hours, though it does not specify that this timing applies equally to every model or test configuration.

At a glance
announcementWhen: Announced in the Hugging Face report; t…
The developmentHugging Face has launched a leaderboard for scalable evaluation of open-source multilingual text-to-speech and voice-cloning models.

Faster Comparisons Across Open TTS

The leaderboard addresses a practical problem for researchers and developers: new speech models are appearing faster than human-vote rankings can test them. Hugging Face’s report says its Hub had more than 8,000 TTS models as of September 30, 2026, while evaluation remained fragmented. A shared set of measurements can help users identify candidates for particular needs, such as accurate multilingual speech, quick response in an interactive application, or preservation of a reference speaker’s voice.

The new system also gives open models more room in comparison tools. The report says that, as of September 30, 2026, 16 of 92 models on Artificial Analysis were open-weight, and describes a similar skew on Voice Arena. It attributes the imbalance in part to operational differences: arena operators can add API services with an API key, while open models must be hosted and served. The leaderboard’s focus on open models may make comparisons more accessible, but its metrics alone cannot establish which speech sounds best to listeners.

Amazon

AI voice recorder with transcription

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How TTS Rankings Have Worked

Existing arena-style TTS rankings generally show people outputs from two systems and ask them to choose a preferred sample. Votes are then used to rank models, often with an Elo score calculated through a Bradley–Terry approach, according to the Hugging Face report. Human preference measures such as MOS or MUSHRA remain important because listeners can judge qualities that automated transcription and timing scores do not capture.

The Open TTS Leaderboard is intended as a complement, not a substitute. Its WER and CER scores provide a proxy for intelligibility, while speaker similarity estimates how closely generated speech resembles a reference voice. Its Pareto plots let users view trade-offs among scores, speed and model size. The report also says Chinese, Japanese and Korean results use CER because they are character-based; the reported average across languages is a macro-average, and Seed TTS Eval has audio only for English and Chinese.

“The Open TTS Leaderboard does not replace human preference ranking.”

— Hugging Face, in its Open TTS Leaderboard report

Limits of the Published Scores

The listed metrics do not measure naturalness, expressiveness or overall listener preference directly, and the report cautions that automated scores cannot settle which output people prefer. The listening feature may add preference data over time, but the report does not give a vote threshold, schedule for adding votes to rankings, or details on how those votes would be weighted.

The source describes the H200 and CPU test settings but does not provide full operational details for every model, including hardware-specific results or the degree to which test conditions are identical across systems. The report’s model rankings and counts are snapshots, with the cited Hub and Artificial Analysis figures dated September 30, 2026. It is not clear how often the leaderboard will be refreshed, which additional datasets may be added, or how community feedback will change its evaluation design.

Community Feedback and New Votes

Hugging Face says users can listen to model outputs and submit feedback through the leaderboard; it asks participants to log in with a Hugging Face account to help reduce spam and bot votes. The report says that, as community votes accumulate, the team may include this information on the leaderboard. It does not announce a specific date for that change or a timetable for later evaluation updates.

For now, readers can use the automated tables and listening tools together: the scores help narrow comparisons by language, speed or voice similarity, while direct listening offers a way to judge qualities the metrics leave out. Any future rankings based on community feedback will depend on how that input is gathered and reported.

Key Questions

What does the Open TTS Leaderboard measure?

It reports WER and CER as proxies for intelligibility, inference speed using inverse real-time factor and TTFA, and speaker similarity based on WavLM embeddings when comparing voice-cloning outputs with a reference.

Does the leaderboard replace human preference tests?

No. Hugging Face says the automated measures do not directly assess naturalness, expressiveness or listener preference. The “Listen” tab lets users compare samples and provide feedback.

Which models lead the reported rankings?

For the reported English WER average, Hugging Face names Kokoro-82M, supertonic-3 and s2-pro among the leading models. It identifies OmniVoice, s2-pro and Fun-CosyVoice3-0.5B-2512 as strong multilingual models; rankings depend on the selected language and evaluation view.

How does the leaderboard assess streaming?

It ranks models by time-to-first-audio, measuring how long it takes for playable audio to arrive. For streaming models, the report describes this as the wait for the first audio chunk; results are tested on an H200 GPU and CPU.

Source: rss

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Claude, Change The “Add To Cart” Button To Blue

An upcoming update will change the ‘Add to Cart’ button color to blue, sparking increased interest in UI customization among online retailers.

Gemini Omni 1.1 Flash

Gemini has launched Omni 1.1 Flash, a firmware update aimed at improving device performance and security, with details confirmed by official sources.

Top Strategies For Reporting AI Model Misalignment In Practice

OpenAI releases a new public framework outlining how it will identify, evaluate, and disclose AI model misbehavior, amid increasing safety concerns.

Can AI Revolutionize The Way We Advertise?

OpenAI releases a statement on how artificial intelligence could reshape advertising, signaling potential future tools and strategies, though specifics remain unconfirmed.