Kokoro-82M: an Apache-2.0 TTS trained for a thousand dollars
kokoro-82m
An open weight TTS model from hexgrad at 82M parameters under Apache-2.0, repository hexgrad/kokoro (9,061 stars, verified through the GitHub API on 2026-09-29). We list it as the cost floor and the licence floor of the audio domain: what deserves to be remembered is not a leaderboard position but the fact that it published the full arithmetic of what one TTS deployment costs. Total training cost was 1000 dollars, being 500 A100 80GB GPU hours for each of v0.19 and v1.0; v0.19 (2024-12-25) trained on under 100 hours of audio with one language and ten voices, while v1.0 (2025-01-27) used a few hundred hours across eight languages and 54 voices. On price, two independent third party sources point at the same order of magnitude: ArtificialAnalysis records 65 cents per million characters on Replicate and DeepInfra lists 80 cents per million characters, which works out to roughly three to five cents per finished hour of audio. That price line is the reference floor for every comparison we make in the audio domain: whatever costs more than it is buying voice cloning, dialect coverage, emotion control, streaming latency or compliance backing, not the ability to be listened to at all. The technical form is a decoder-only StyleTTS 2 plus ISTFTNet stack with no diffusion and an unpublished encoder, G2P through the author maintained misaki library, espeak-ng on the system side and 24kHz output; input text accepts inline phoneme overrides such as [Kokoro](/kˈOkəɹO/), the same idea as the pinyin and CMU phoneme level pronunciation correction in CosyVoice 3, turning misreading from a probability problem into declarative configuration, except the scope here is a word rather than a whole prosody. The decoder-only shape also decides what it cannot do: there is no zero shot voice cloning, the 54 voices are trained in and a reference clip will not grow a new one, which is a division of labour rather than a defect. The compliance disclosure is close to unique in TTS: the model card states that only permissively licensed or copyright free audio was used, lists public domain audio, Apache and MIT licensed audio and synthetic audio generated by closed source TTS models as separate classes, explicitly excludes synthetic audio from open TTS models and custom voice cloning, publishes a CC BY attribution table, and publishes the model SHA256 so you can verify the weights you downloaded are the same ones. Boundaries: no voice cloning, so reference driven use cases are out; 24kHz is enough for broadcast and reading but not for music grade masters; the eight languages and 54 voices have not grown since v1.0 and Mandarin usability needs your own audition; the project is close to dormant with the last push on 2025-08-06, so bugs are yours to fix; and its fame has grown a phishing surface, with the model card naming lookalike domains containing kokoro as unrelated to the project, the only real entries being hexgrad/Kokoro-82M on Hugging Face and hexgrad/kokoro on GitHub, so never pay a search result that calls itself the Kokoro official site. Graded B (confirmed): what is confirmed is the cost and market price chain with independent third party sources, not audio quality.
- CONFIDENCE
- Confirmed
- Two or more independent sources, or reproduced by our harness
- KEY METRIC
- API 市场价(每百万字符,2025-04 模型卡口径)
- Confirmed · 2025-04
- MATURITY
- Product
- research → demo → product → production
Our takeWe grade it B (confirmed), and it is worth being precise about which claim is confirmed. It is not sound quality, it is cost. The model card states the April 2025 market rate directly: Kokoro served over an API costs under one dollar per million characters of text input, or under 0.06 dollar per hour of audio output, with roughly 1000 characters of input producing about one minute of output. It cites two independent sources, ArtificialAnalysis recording Replicate at 65 cents per million characters and DeepInfra listing
hexgrad/Kokoro-82Mat 80 cents per million characters. Two independent third party price points landing in the same band is enough for grade B under our contract. The repository stands at 9,061 stars under Apache-2.0, verified by us through the GitHub API on 2026-09-29. On quality we pass the numbers through without recomputing them and we have run no blind listening test.Its role in the audio ladder is different from the other rows: it is the cost floor and the licence floor of this domain. The most common mistake in TTS selection is budgeting every use case against a flagship metric, when a large share of real demand (notification playback, accessibility reading, internal tooling, edge devices) needs no cloning at all and only needs cheap, commercially usable, self-hostable and reliable. Kokoro pins that line with 82 million parameters and about 1000 dollar of training cost: 1000 A100-80GB GPU hours, 500 for v0.19 and 500 for v1.0, at roughly one dollar per hour. Once that floor exists, evaluating a flagship such as Eleven v4 or CosyVoice 3 becomes a question about what the extra money actually buys.
The second reason is auditable training data provenance, which is close to unique in TTS. The model card states that training used only permissive or non-copyrighted audio plus IPA phoneme labels, and it enumerates the categories: public domain audio, audio licensed under Apache or MIT, and synthetic audio generated by closed TTS models from large providers, citing the AI policy guidance of the United States Copyright Office. It explicitly excludes synthetic audio from open TTS models and any custom voice clones. It also publishes a CC BY attribution table naming Koniwa
tnc(under one hour, CC BY 3.0) and SIWIS (under eleven hours, CC BY 4.0), both added to the training set after 22 November 2024. For enterprise procurement this disclosure is worth far more than a sentence about regulatory compliance, because it lets a buyer assess the risk instead of being told the risk has been assessed.Three boundaries. First, it does not clone voices: the architecture is decoder-only StyleTTS 2 with ISTFTNet (no diffusion, no encoder release), and v1.0 offers 8 languages across 54 fixed voices. You choose a voice, you cannot hand it a reference recording and get a new one. That is a division of labour with CosyVoice 3 and Eleven v4 rather than a capability gap. Second, development has slowed: the last push to the repository was 2025-08-06, so the honest framing is a mature small model that has essentially stopped iterating, not a project still growing. Third, its fame has grown a phishing surface: the model card carries an explicit warning that sites with kokoro in the root domain such as
kokorottsai_comandkokorotts_netare not owned by and not affiliated with the model or its author, with archive snapshots attached. That has practical value for our readers: self-host from the HF weights andpip install kokoro, and do not pay a site that a search engine labelled as the official Kokoro page.
The problem it solves: what a TTS deployment costs when the model is free
Kokoro is the open-weight TTS model published by hexgrad, 82 million parameters, Apache-2.0, repository hexgrad/kokoro (9,061 stars, last push 2025-08-06, verified by us through the GitHub API). The model card states its case without inflation: despite the lightweight architecture it delivers quality comparable to much larger models while being significantly faster and more cost efficient, and because the weights are Apache licensed it can be deployed anywhere from production environments to personal projects. v1.0 shipped on 2025-01-27 with a few hundred hours of training data and 8 languages across 54 voices; the preceding v0.19 (2024-12-25) had under 100 hours, 1 language and 10 voices.
What it is best remembered for is not a leaderboard position but for publishing the cost structure of a TTS model:
| Item | v0.19 | v1.0 | Total |
|---|---|---|---|
| A100 80GB GPU hours | 500 | 500 | 1000 |
| Average hourly rate | $0.80/h | $1.20/h | about $1/h |
| USD | $400 | $600 | $1000 |
| Training data | <100 hours | a few hundred hours | - |
| Languages / voices | 1 / 10 | 8 / 54 | - |
A thousand dollars producing a model that numerous commercial APIs then served matters because it moved training a usable TTS system from a large-lab activity to something an individual could attempt. The acknowledgements in the model card record how it became visible: thanks to @Pendrokar for adding Kokoro as a contender in the TTS Spaces Arena, thanks to everyone who contributed synthetic training data, and thanks to the compute sponsors. It was trained by @rzvzn on Discord, on an architecture by Li and colleagues (StyleTTS 2), with yl4579/StyleTTS2-LJSpeech as the base model.
Market price: two independent sources, one band
The April 2025 note in the model card gives the rate directly: Kokoro served over an API costs under one dollar per million characters of text input, which is under 0.06 dollar per hour of audio output, with about 1000 characters of input producing one minute of output on average. The two cited sources are independent:
- ArtificialAnalysis, recording Replicate at 65 cents per million characters.
- DeepInfra, listing
hexgrad/Kokoro-82Mat 80 cents per million characters.
Converted into operational language, one hour of finished audio costs roughly three to five cents in speech generation. That is the reference floor we use for comparison inside the audio domain: whatever costs more than this is buying voice cloning, dialect coverage, emotion control, streaming latency or compliance backing, not basic intelligibility.
Architecture: decoder only, which is why it does not clone
The architecture is StyleTTS 2 with ISTFTNet and it is decoder only: no diffusion and no encoder release. Grapheme to phoneme conversion runs through misaki, a library the author maintains, with espeak-ng as a system dependency and 24 kHz output. The minimum runnable path is a few lines:
pip install "kokoro>=0.9.2" soundfile
apt-get -qq -y install espeak-ng
from kokoro import KPipeline
import soundfile as sf
pipeline = KPipeline(lang_code="a") # language code, a = American English
text = "Kokoro is an open-weight TTS model with 82 million parameters."
generator = pipeline(text, voice="af_heart") # one of the 54 fixed voices
for i, (gs, ps, audio) in enumerate(generator):
sf.write(f"{i}.wav", audio, 24000)
One design detail is easy to miss: the input text accepts inline phoneme overrides, written [Kokoro](/kˈOkəɹO/), so the reading of a specific word can be pinned without touching the model. That is the same idea as the pinyin and CMU phoneme inpainting in CosyVoice 3, converting mispronunciation from a probability problem into a declaration, except the scope here is a word rather than a span of prosody. The decoder-only design is also what determines that there is no zero-shot voice cloning: the 54 voices are trained in, and a reference recording will not produce a new one. That is a division of labour, not a defect.
Training data: a compliance table you can audit
The provenance disclosure in the model card has close to no equivalent in TTS. It states that Kokoro was trained exclusively on permissive or non-copyrighted audio plus IPA phoneme labels, and it names three categories: public domain audio, audio licensed under Apache or MIT, and synthetic audio generated by closed TTS models from large providers, with a footnote citing the AI policy guidance of the United States Copyright Office. A second footnote excludes synthetic audio from open TTS models and any custom voice clones. The CC BY portion gets its own attribution table:
| Audio data | Duration used | Licence | Added to the training set after |
|---|---|---|---|
Koniwa tnc | <1 hour | CC BY 3.0 | v0.19 / 22 November 2024 |
| SIWIS | <11 hours | CC BY 4.0 | v0.19 / 22 November 2024 |
The model SHA256 is published as well (496dba11...), so a deployer can verify that the downloaded weights are the intended artefact. For a compliance sensitive buyer, these three things together (a categorised data provenance statement, an attribution table and a weight hash) are what the words Apache-2.0 are actually worth.
Boundaries and failure modes
Five of them. No voice cloning: any requirement driven by a reference recording is out of scope, look at CosyVoice 3 for open weight or Eleven v4 for closed. 24 kHz output: enough for playback and reading, not enough for music grade masters, so a 48 kHz pipeline either resamples afterwards or picks another model. Languages and voices are a closed set: 8 languages and 54 voices, not extended after v1.0, and practical quality outside English needs a local listening test. The project has essentially stopped moving: the last push was 2025-08-06, so there will be no new features, which buys stability and means any bug you hit is yours to patch. Phishing sites: the model card names kokorottsai_com and kokorotts_net as unaffiliated with the model and its author, attaches archive.ph snapshots, and states that any site with kokoro in the root domain implying an affiliation is a red flag. There are only two real entry points, hexgrad/Kokoro-82M on Hugging Face and hexgrad/kokoro on GitHub.