Hugging Face launched the Open TTS Leaderboard on 30 September 2026. It ranks 31 open text-to-speech models on intelligibility, speed and voice-cloning similarity, scored by software instead of human votes. On English, the top model is Kokoro-82M, an Apache 2.0 model with 82 million parameters, ahead of models 60 times its size.
We read the license of every model on the board and timed the two CPU entries on our own machine. The rankings hold up. The license column does not. Switch to multilingual or voice cloning and the #1 and #2 models are released for research and non-commercial use only. The first model you can put in a paid product is Qwen3-TTS at #3. The leaderboard's own license column says "custom" for 8 of the 29 English models, which hides exactly that.
What the Open TTS Leaderboard measures
The launch post by Eric Bezzam, Steven Zheng, Eustache Le Bihan and mrfakename makes the case plainly: there are more than 8,000 TTS models on the Hub, and arenas cannot keep up. Human-vote leaderboards take weeks per model, and open weights are underrepresented on them. Only 16 of the 92 models on Artificial Analysis are open-weights, per the post.
So this board scores three things automatically, in a couple of hours per model:
- WER (word error rate). The generated speech is transcribed by Qwen3-ASR 1.7B and compared with the input text. Lower means the words came out right. Chinese, Japanese and Korean use character error rate.
- RTFx (speed). Seconds of audio produced per second of compute, on one H200 GPU, at batch size 32 where a model supports batching and 1 where it does not.
- SIM (speaker similarity). For voice cloning only: how close the generated voice is to the reference clip, from WavLM speaker embeddings.
Test sentences come from Seed-TTS eval (English and Chinese) and CV3-Eval (nine languages). A Streaming tab adds time to first audio (TTFA) at batch size 1 on GPU and, for four entries, on CPU. A Listen tab plays the actual clips behind the numbers.
None of this measures how natural a voice sounds. The authors say so directly: the board does not replace human preference. It tells you whether a model says the words correctly, how fast it runs, and how well it copies a voice.
The English ranking: an 82M model on top
Here are the top 10 on the default view (English, default voice, both datasets averaged), with the license class we found by reading each license text.
| Rank | Model | WER | RTFx (H200) | Size | Commercial use |
|---|---|---|---|---|---|
| 1 | Kokoro-82M | 1.46 | 71.15 | 0.08B | Yes (Apache 2.0) |
| 2 | Supertonic 3 | 1.54 | 47.32 | 0.1B | Yes, with use restrictions (Open RAIL-M) |
| 3 | Fish Audio S2 Pro | 1.58 | 0.4 | 4.95B | No, separate license needed |
| 4 | OmniVoice | 1.61 | 37.56 | 0.81B | No (CC-BY-NC weights) |
| 5 | pocket-tts | 1.62 | 9.09 | 0.1B | Yes (CC-BY-4.0, credit Kyutai) |
| 6 | Fun-CosyVoice3 0.5B | 1.68 | 5.88 | 0.11B | Yes (Apache 2.0) |
| 7 | Chatterbox | 1.70 | 2.07 | 0.8B | Yes (MIT) |
| 8 | Qwen3-TTS 1.7B CustomVoice | 1.71 | 13.32 | 1.92B | Yes (Apache 2.0) |
| 9 | Breeze-TTS-2 | 1.72 | 3.32 | 3.47B | No, separate license needed |
| 10 | VibeVoice-Realtime 0.5B | 1.82 | 2.71 | 1.36B | Yes (MIT) |
Two things stand out. First, the spread is tiny: #1 to #10 runs from 1.46 to 1.82 WER, meaning roughly one to two words wrong per hundred for every model in the top 10. At that range, the ASR judge's own mistakes are part of the score, so treat the top five as a group, not a podium.
Second, size does not buy accuracy here. Kokoro has 82 million parameters. Fish Audio S2 Pro has 4.95 billion, about 60 times more, and scores 1.58 against Kokoro's 1.46. On the same H200, Kokoro produced 71.15 seconds of audio per second of compute; S2 Pro produced 0.4, slower than real time. Kokoro is also not new: its Hub repo was last updated in April 2025, and it was downloaded about 11.5 million times in the past month.
What Kokoro lacks is voice cloning. It speaks in its built-in voices only, and it covers 8 languages. Supertonic 3 at #2 covers 31, which we covered at launch.

Our check: what the license column hides
The leaderboard has a License column. For 8 of the 29 English models it says "custom". We opened each license file on the Hub, plus the model cards and one GitHub license, and sorted all 29 into what you can actually do with them.
| What the license allows | Models | Count |
|---|---|---|
| Commercial use, attribution at most (Apache 2.0, MIT, CC-BY-4.0, NVIDIA Open Model License) | Kokoro, pocket-tts, Qwen3-TTS, Chatterbox, CosyVoice3, VoxCPM2, MOSS-TTS, AuK, VibeVoice, Magpie and others | 18 |
| Commercial, with use-based restrictions (Open RAIL-M) | Supertonic 3 | 1 |
| Commercial below a size threshold | Higgs TTS 2 (100,000 annual active users), IndexTTS 2.5 (100M monthly users or RMB 1B revenue), NeuTTS Nano and its GGUF build ($5M annual revenue) | 4 |
| Research and non-commercial only | Fish Audio S2 Pro, OmniVoice, Breeze-TTS-2, SeamlessM4T v2, Voxtral 4B TTS | 5 |
| Non-commercial, plus a free grant for creators' own content | Higgs TTS 3 | 1 |
So 19 of 29 are clean for commercial use, 4 are clean for almost everyone, and 6 need a conversation with the vendor before they go anywhere near revenue. Three of those six sit in the English top 10.
The "custom" label covers all three kinds. The Fish Audio Research License says any commercial use needs a separate written agreement. The NeuTTS Open License allows commercial use until your company reaches $5 million in annual revenue. NVIDIA's Magpie is labelled "custom" too, and its license states that the models are commercially usable. One label, three different answers.
We found one mismatch the other way. The leaderboard lists Fun-CosyVoice3 as CC-BY-4.0; its model card says Apache 2.0. Both allow commercial use, but it shows the column is typed by hand.
Multilingual and voice cloning: the top two are non-commercial
Tick all nine languages and only 9 of 31 models have results for every one, so only those 9 get ranked. The top of that list:
| View | #1 | #2 | #3 |
|---|---|---|---|
| 9 languages, default voice (avg WER) | OmniVoice 3.25 (non-commercial) | Fish Audio S2 Pro 3.68 (non-commercial) | Qwen3-TTS 1.7B CustomVoice 4.11 (Apache 2.0) |
| 9 languages, voice cloning (avg WER / SIM) | OmniVoice 3.52 / 74.77 (non-commercial) | Higgs TTS 3 3.74 / 69.08 (non-commercial, creator grant) | Qwen3-TTS 1.7B Base 3.92 / 72.92 (Apache 2.0) |
| English, voice cloning (WER / SIM) | Higgs TTS 3 1.59 / 64.69 (non-commercial, creator grant) | OmniVoice 1.69 / 71.98 (non-commercial) | Qwen3-TTS 1.7B Base 1.81 / 69.33 (Apache 2.0) |
In every one of these views, the first model licensed for commercial use is Qwen3-TTS at #3. OmniVoice explains its restriction in its README: the code is Apache 2.0, but the weights are CC-BY-NC because of the training data, including the Emilia dataset. We covered OmniVoice at launch in April; the license was the fine print then and it is the headline now.
Higgs TTS 3 is the interesting exception. Its license is research and non-commercial, but section II-A adds a free Creator Use Grant. You may generate audio for podcasts, videos, audiobooks and social posts, and monetize them on channels you own, including ad-supported and sponsored ones. The one condition is a visible credit in the audio or in the description, such as "This audio was created with Boson AI's Higgs Audio." What the grant does not cover is hosting the model for others, reselling it, or building it into an app or a dubbing service. For a solo creator voicing their own channel, that makes Higgs TTS 3 usable. For anyone building a product, it does not.
If voice match matters most, look past the WER rank. On the nine-language cloning view, VoxCPM2 has the highest speaker similarity of the seven ranked models, 75.19, under Apache 2.0, at #5 on WER. On English cloning, Tencent's AuK scores the highest similarity of all, 75.08, under MIT, but it covers only English and Chinese.

Our CPU test: pocket-tts starts talking 5 times sooner
Only four entries have CPU results, and they are the ones creators can run on a laptop or a cheap server. We reproduced the leaderboard's streaming method for the two distinct models: batch size 1, default voice, 3 warm-up sentences, then 47 timed sentences from CV3-Eval's English set. Our machine is a 6-vCPU container on an AMD Ryzen 7 9700X, with PyTorch 2.14.1 CPU.
| Setup | pocket-tts TTFA | pocket-tts RTFx | Kokoro TTFA | Kokoro RTFx |
|---|---|---|---|---|
| Leaderboard CPU (Hugging Face cpu-upgrade) | 272.7 ms | 1.19 | 913.4 ms | 6.21 |
| Ours, 6 threads | 69.3 ms | 5.02 | 366.5 ms | 11.36 |
| Ours, 2 threads | 71.1 ms | 4.85 | 701.5 ms | 6.14 |
The order held. pocket-tts started speaking in a median 69 ms against 367 ms for Kokoro on 6 threads, a 5.3x gap, wider than the leaderboard's 3.3x. Kokoro produced more audio per second, 11.36 against 5.02, because it renders a whole sentence at once while pocket-tts streams it in small chunks.
Thread count split them further. Cut to 2 threads and pocket-tts barely moved: 71.1 ms and 4.85x real time. Kyutai says it uses only 2 CPU cores and runs about 6x real time on a MacBook Air M4, and our result is close to that. Kokoro, on 2 threads, nearly doubled its wait to 701.5 ms and fell to 6.14x.
Our absolute numbers are roughly 2 to 4 times better than the leaderboard's CPU column. That is hardware, not a flaw in either measurement: a desktop-class Ryzen is faster than a shared cloud vCPU. Read the leaderboard's CPU numbers as a floor for your own machine, and trust the ratios more than the milliseconds.
One practical catch. Without a Hugging Face login, the pip package downloaded pocket-tts-without-voice-cloning. The voice-cloning weights sit in a gated repo where you must accept Kyutai's terms first, and those terms prohibit impersonation and cloning a voice without consent.

Speed depends on the serving code, not only the model
The Streaming tab shows how much the software around a model matters. Qwen3-TTS 1.7B CustomVoice reaches first audio in 18.3 ms on the H200 when served with Moondream's Photon engine, and in 309.9 ms with the faster-qwen3-tts package. Same weights, 17 times apart. The Photon rows were added in the 30 September update. We measured the same effect on a different leaderboard in our Qwen3-TTS serving test.
Two more rows to read carefully. Supertonic 3, #2 on English WER, takes 1,918.6 ms to first audio because it has no streaming mode; it is fast in batches (47.32 RTFx) and slow for a single live sentence (3.53). And Fish Audio S2 Pro shows 0.4 RTFx on the main table but 5.1 on the Streaming tab, a 13x difference for the same model on the same GPU that the leaderboard does not explain.
How to pick an open TTS model from this board
- Decide whether the audio earns money. Paid products, client work and apps rule out the 5 non-commercial models. Your own monetized channel adds Higgs TTS 3, with the credit line.
- English voiceover on a CPU: Kokoro-82M. Apache 2.0, #1 on English WER, about 11 times real time on our 6 threads. Install with
pip install kokoro soundfile; we used theaf_heartvoice. - A live voice or agent on a CPU: pocket-tts. The fastest first audio on CPU, and it stays fast on 2 cores. CC-BY-4.0, so credit Kyutai. Try it with
uvx pocket-tts generate. - Many languages, commercial: Qwen3-TTS 1.7B, CustomVoice for built-in voices or Base for cloning. It is the top Apache 2.0 model in every multilingual view, and the Photon engine gives it the fastest start on GPU.
- Closest voice match, commercial: VoxCPM2 across languages, AuK for English and Chinese.
- Listen before you commit. Open the Listen tab, filter to your language, and play the stored clips from your shortlist. WER tells you the words are right, not that the voice suits your content.
For the other half of the market, closed models ranked by human votes, see where Eleven v4 lands on price. For the tools that wrap these models in an app, see our open-source voice cloning comparison.
What the leaderboard does not tell you yet
- Whether a voice sounds good. WER and SIM are proxies. Votes from the Listen tab may be added later, the authors say.
- Most languages. 29 models have English results; only 9 have all nine languages. Many models are ranked on one or two.
- Reproducibility, for now. The evaluation scripts are "planned" to be open-sourced, and the results dataset (
hf-audio/tts_leaderboard_results) returned 401 to us. We read the numbers through the Space's own API instead. - Stability. The board is versioned, with a results-version dropdown. The version we used is dated 30-09-2026; later updates will move rows.
How we checked
We pulled every table (English, English cloning, nine languages, nine-language cloning, GPU and CPU streaming) through the Space's public API for results version 30-09-2026. For licenses, we read the license field of all 32 model cards on the board, then the full license texts behind the 10 entries without a standard open license, including the NeuTTS license on GitHub because its Hub repo is gated. For speed, we ran pocket-tts 3.3.0 and Kokoro 0.9.4 on the setup above, 3 warm-up and 47 timed sentences each, at 6 and 2 threads. We did not run a listening test and did not re-measure WER. Scripts and raw results are kept with this article.
Frequently asked questions
What is the Open TTS Leaderboard?
A Hugging Face Space, launched on 30 September 2026, that ranks open-source text-to-speech models on word error rate, speed and voice-cloning similarity using automatic metrics instead of human votes. It covers 31 models across up to nine languages.
What is the best open-source TTS model in 2026?
On English intelligibility, Kokoro-82M ranks #1 with 1.46 WER, and it is Apache 2.0. Across nine languages OmniVoice ranks #1, but its weights are non-commercial, so the best commercial pick is Qwen3-TTS 1.7B at #3.
Can I use OmniVoice or Fish Audio S2 Pro commercially?
Not under their public licenses. OmniVoice weights are CC-BY-NC, and the Fish Audio Research License requires a separate written license for any commercial use. Personal and research use is free.
Can creators monetize audio made with Higgs TTS 3?
Yes, on channels you own or control, under the Creator Use Grant in its license, as long as you credit Boson AI's Higgs Audio visibly in the audio or the description. You cannot host it for others or build it into a product without a commercial license.
Which open TTS model runs fastest on a CPU?
For time to first audio, pocket-tts: a median 69 ms on our 6-thread test and 272.7 ms on the leaderboard's CPU. For total throughput, Kokoro-82M produced about 11 seconds of audio per second on our 6 threads.
Does the Open TTS Leaderboard measure voice quality?
No. It measures intelligibility (whether the words come out right), speed and speaker similarity. Naturalness still needs human listening, which is what the Listen tab and arenas like TTS Arena are for.