Open TTS Leaderboard Benchmarks Multilingual Speech, Speed and Voice Cloning
Open TTS Leaderboard benchmarks 8,000 plus speech models with multilingual accuracy, voice similarity and H200 speed metrics in hours, not weeks. for developers
Summary
Released September 30, 2026, the Open TTS Leaderboard addresses fragmented evaluation across more than 8,000 TTS models on the Hugging Face Hub. Human arenas TTS Arena v2, Artificial Analysis and Voice Arena rank pairwise votes with Elo, typically using Bradley Terry, but take weeks, vary with voters and underrepresent open weights: only 16 of Artificial Analysis’s 92 models were open weights on September 30. API models need only a key, while open models require arena hosting and commercial vendors have stronger placement incentives. Objective testing cuts evaluation to hours: Qwen3 ASR, the top open-source Open ASR Leaderboard model, measures WER and CER intelligibility; an H200 measures batched inverse real-time factor, H200 and CPU runs measure batch size 1 TTFA, and WavLM embedding cosine similarity against a reference clip tests cloned speaker identity. These proxies do not measure naturalness, expressiveness or preference.
The default English ranking macro-averages Seed TTS Eval and zero-shot CV3 Eval: hexgrad/Kokoro-82M, Supertone/supertonic-3 and fishaudio/s2-pro lead. Multilingual leaders are k2-fsa/OmniVoice, fishaudio/s2-pro and FunAudioLLM/Fun-CosyVoice3-0.5B-2512. Seed covers English and Chinese, other languages use CV3, Chinese, Japanese and Korean use CER, and the cross-language score is a macro-average. Voice cloning adds SIM and trade-off views against throughput and size; bosonai/higgs-tts-3-4b and openbmb/VoxCPM2 improve WER with reference audio. Streaming ranks median TTFA on H200, plus CPU for a small, growing set, using 50 English CV3-Eval prompts, default voices, batch size 1 and three discarded warmups; kyutai/pocket-tts excels on both. The Listen tab exposes outputs and accepts votes from logged-in Hugging Face accounts. Evaluation scripts will soon be open sourced for GitHub Issues and pull requests.
Positives
- Objective metrics reduce each model evaluation from several weeks of vote collection to several hours.
- hexgrad/Kokoro-82M, Supertone/supertonic-3 and fishaudio/s2-pro lead the combined English WER ranking.
- k2-fsa/OmniVoice, fishaudio/s2-pro and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 show strong multilingual performance.
- kyutai/pocket-tts delivers strong streaming performance on both H200 GPU and CPU.
- Logged-in Hugging Face users can compare underlying audio and submit votes through the Listen tab.
- Evaluation scripts will soon be open sourced for community changes through GitHub Issues and pull requests.
Risks & concerns
- Only 16 of Artificial Analysis’s 92 models were open weights on September 30, 2026.
- WER, CER and speaker similarity cannot directly measure naturalness, expressiveness or listener preference.
- Arena rankings can shift because the same voters and evaluation criteria cannot remain consistent over time.
- Open models require hosting and serving by arena operators, while API models generally need only an API key.
- CPU streaming results cover only a small set of models, although that selection is growing.
- Seed TTS Eval provides audio only for English and Chinese, leaving CV3 Eval to supply other language scores.