Skip to content
←Back to Applications

APPLICATION

Speech & AudioMediaTop

Kokoro-82M: an Apache-2.0 TTS trained for a thousand dollars

kokoro-82m

An open weight TTS model from hexgrad at 82M parameters under Apache-2.0, repository hexgrad/kokoro (9,061 stars, verified through the GitHub API on 2026-09-29). We list it as the cost floor and the licence floor of the audio domain: what deserves to be remembered is not a leaderboard position but the fact that it published the full arithmetic of what one TTS deployment costs. Total training cost was 1000 dollars, being 500 A100 80GB GPU hours for each of v0.19 and v1.0; v0.19 (2024-12-25) trained on under 100 hours of audio with one language and ten voices, while v1.0 (2025-01-27) used a few hundred hours across eight languages and 54 voices. On price, two independent third party sources point at the same order of magnitude: ArtificialAnalysis records 65 cents per million characters on Replicate and DeepInfra lists 80 cents per million characters, which works out to roughly three to five cents per finished hour of audio. That price line is the reference floor for every comparison we make in the audio domain: whatever costs more than it is buying voice cloning, dialect coverage, emotion control, streaming latency or compliance backing, not the ability to be listened to at all. The technical form is a decoder-only StyleTTS 2 plus ISTFTNet stack with no diffusion and an unpublished encoder, G2P through the author maintained misaki library, espeak-ng on the system side and 24kHz output; input text accepts inline phoneme overrides such as [Kokoro](/kˈOkəɹO/), the same idea as the pinyin and CMU phoneme level pronunciation correction in CosyVoice 3, turning misreading from a probability problem into declarative configuration, except the scope here is a word rather than a whole prosody. The decoder-only shape also decides what it cannot do: there is no zero shot voice cloning, the 54 voices are trained in and a reference clip will not grow a new one, which is a division of labour rather than a defect. The compliance disclosure is close to unique in TTS: the model card states that only permissively licensed or copyright free audio was used, lists public domain audio, Apache and MIT licensed audio and synthetic audio generated by closed source TTS models as separate classes, explicitly excludes synthetic audio from open TTS models and custom voice cloning, publishes a CC BY attribution table, and publishes the model SHA256 so you can verify the weights you downloaded are the same ones. Boundaries: no voice cloning, so reference driven use cases are out; 24kHz is enough for broadcast and reading but not for music grade masters; the eight languages and 54 voices have not grown since v1.0 and Mandarin usability needs your own audition; the project is close to dormant with the last push on 2025-08-06, so bugs are yours to fix; and its fame has grown a phishing surface, with the model card naming lookalike domains containing kokoro as unrelated to the project, the only real entries being hexgrad/Kokoro-82M on Hugging Face and hexgrad/kokoro on GitHub, so never pay a search result that calls itself the Kokoro official site. Graded B (confirmed): what is confirmed is the cost and market price chain with independent third party sources, not audio quality.

A
CONFIDENCE
Confirmed
Two or more independent sources, or reproduced by our harness
<$1
KEY METRIC
API 市场价(每百万字符,2025-04 模型卡口径)
Confirmed · 2025-04
MATURITY
Product
research → demo → product → production
Our take

We grade it B (confirmed), and it is worth being precise about which claim is confirmed. It is not sound quality, it is cost. The model card states the April 2025 market rate directly: Kokoro served over an API costs under one dollar per million characters of text input, or under 0.06 dollar per hour of audio output, with roughly 1000 characters of input producing about one minute of output. It cites two independent sources, ArtificialAnalysis recording Replicate at 65 cents per million characters and DeepInfra listing hexgrad/Kokoro-82M at 80 cents per million characters. Two independent third party price points landing in the same band is enough for grade B under our contract. The repository stands at 9,061 stars under Apache-2.0, verified by us through the GitHub API on 2026-09-29. On quality we pass the numbers through without recomputing them and we have run no blind listening test.

Its role in the audio ladder is different from the other rows: it is the cost floor and the licence floor of this domain. The most common mistake in TTS selection is budgeting every use case against a flagship metric, when a large share of real demand (notification playback, accessibility reading, internal tooling, edge devices) needs no cloning at all and only needs cheap, commercially usable, self-hostable and reliable. Kokoro pins that line with 82 million parameters and about 1000 dollar of training cost: 1000 A100-80GB GPU hours, 500 for v0.19 and 500 for v1.0, at roughly one dollar per hour. Once that floor exists, evaluating a flagship such as Eleven v4 or CosyVoice 3 becomes a question about what the extra money actually buys.

The second reason is auditable training data provenance, which is close to unique in TTS. The model card states that training used only permissive or non-copyrighted audio plus IPA phoneme labels, and it enumerates the categories: public domain audio, audio licensed under Apache or MIT, and synthetic audio generated by closed TTS models from large providers, citing the AI policy guidance of the United States Copyright Office. It explicitly excludes synthetic audio from open TTS models and any custom voice clones. It also publishes a CC BY attribution table naming Koniwa tnc (under one hour, CC BY 3.0) and SIWIS (under eleven hours, CC BY 4.0), both added to the training set after 22 November 2024. For enterprise procurement this disclosure is worth far more than a sentence about regulatory compliance, because it lets a buyer assess the risk instead of being told the risk has been assessed.

Three boundaries. First, it does not clone voices: the architecture is decoder-only StyleTTS 2 with ISTFTNet (no diffusion, no encoder release), and v1.0 offers 8 languages across 54 fixed voices. You choose a voice, you cannot hand it a reference recording and get a new one. That is a division of labour with CosyVoice 3 and Eleven v4 rather than a capability gap. Second, development has slowed: the last push to the repository was 2025-08-06, so the honest framing is a mature small model that has essentially stopped iterating, not a project still growing. Third, its fame has grown a phishing surface: the model card carries an explicit warning that sites with kokoro in the root domain such as kokorottsai_com and kokorotts_net are not owned by and not affiliated with the model or its author, with archive snapshots attached. That has practical value for our readers: self-host from the HF weights and pip install kokoro, and do not pay a site that a search engine labelled as the official Kokoro page.

Text-to-SpeechOpen WeightsApache-2.0Cost Efficiency

The problem it solves: what a TTS deployment costs when the model is free

Kokoro is the open-weight TTS model published by hexgrad, 82 million parameters, Apache-2.0, repository hexgrad/kokoro (9,061 stars, last push 2025-08-06, verified by us through the GitHub API). The model card states its case without inflation: despite the lightweight architecture it delivers quality comparable to much larger models while being significantly faster and more cost efficient, and because the weights are Apache licensed it can be deployed anywhere from production environments to personal projects. v1.0 shipped on 2025-01-27 with a few hundred hours of training data and 8 languages across 54 voices; the preceding v0.19 (2024-12-25) had under 100 hours, 1 language and 10 voices.

What it is best remembered for is not a leaderboard position but for publishing the cost structure of a TTS model:

Itemv0.19v1.0Total
A100 80GB GPU hours5005001000
Average hourly rate$0.80/h$1.20/habout $1/h
USD$400$600$1000
Training data<100 hoursa few hundred hours-
Languages / voices1 / 108 / 54-

A thousand dollars producing a model that numerous commercial APIs then served matters because it moved training a usable TTS system from a large-lab activity to something an individual could attempt. The acknowledgements in the model card record how it became visible: thanks to @Pendrokar for adding Kokoro as a contender in the TTS Spaces Arena, thanks to everyone who contributed synthetic training data, and thanks to the compute sponsors. It was trained by @rzvzn on Discord, on an architecture by Li and colleagues (StyleTTS 2), with yl4579/StyleTTS2-LJSpeech as the base model.

Market price: two independent sources, one band

The April 2025 note in the model card gives the rate directly: Kokoro served over an API costs under one dollar per million characters of text input, which is under 0.06 dollar per hour of audio output, with about 1000 characters of input producing one minute of output on average. The two cited sources are independent:

  • ArtificialAnalysis, recording Replicate at 65 cents per million characters.
  • DeepInfra, listing hexgrad/Kokoro-82M at 80 cents per million characters.

Converted into operational language, one hour of finished audio costs roughly three to five cents in speech generation. That is the reference floor we use for comparison inside the audio domain: whatever costs more than this is buying voice cloning, dialect coverage, emotion control, streaming latency or compliance backing, not basic intelligibility.

Architecture: decoder only, which is why it does not clone

The architecture is StyleTTS 2 with ISTFTNet and it is decoder only: no diffusion and no encoder release. Grapheme to phoneme conversion runs through misaki, a library the author maintains, with espeak-ng as a system dependency and 24 kHz output. The minimum runnable path is a few lines:

pip install "kokoro>=0.9.2" soundfile
apt-get -qq -y install espeak-ng

from kokoro import KPipeline
import soundfile as sf

pipeline = KPipeline(lang_code="a")        # language code, a = American English
text = "Kokoro is an open-weight TTS model with 82 million parameters."
generator = pipeline(text, voice="af_heart")   # one of the 54 fixed voices
for i, (gs, ps, audio) in enumerate(generator):
    sf.write(f"{i}.wav", audio, 24000)

One design detail is easy to miss: the input text accepts inline phoneme overrides, written [Kokoro](/kˈOkəɹO/), so the reading of a specific word can be pinned without touching the model. That is the same idea as the pinyin and CMU phoneme inpainting in CosyVoice 3, converting mispronunciation from a probability problem into a declaration, except the scope here is a word rather than a span of prosody. The decoder-only design is also what determines that there is no zero-shot voice cloning: the 54 voices are trained in, and a reference recording will not produce a new one. That is a division of labour, not a defect.

Training data: a compliance table you can audit

The provenance disclosure in the model card has close to no equivalent in TTS. It states that Kokoro was trained exclusively on permissive or non-copyrighted audio plus IPA phoneme labels, and it names three categories: public domain audio, audio licensed under Apache or MIT, and synthetic audio generated by closed TTS models from large providers, with a footnote citing the AI policy guidance of the United States Copyright Office. A second footnote excludes synthetic audio from open TTS models and any custom voice clones. The CC BY portion gets its own attribution table:

Audio dataDuration usedLicenceAdded to the training set after
Koniwa tnc<1 hourCC BY 3.0v0.19 / 22 November 2024
SIWIS<11 hoursCC BY 4.0v0.19 / 22 November 2024

The model SHA256 is published as well (496dba11...), so a deployer can verify that the downloaded weights are the intended artefact. For a compliance sensitive buyer, these three things together (a categorised data provenance statement, an attribution table and a weight hash) are what the words Apache-2.0 are actually worth.

Boundaries and failure modes

Five of them. No voice cloning: any requirement driven by a reference recording is out of scope, look at CosyVoice 3 for open weight or Eleven v4 for closed. 24 kHz output: enough for playback and reading, not enough for music grade masters, so a 48 kHz pipeline either resamples afterwards or picks another model. Languages and voices are a closed set: 8 languages and 54 voices, not extended after v1.0, and practical quality outside English needs a local listening test. The project has essentially stopped moving: the last push was 2025-08-06, so there will be no new features, which buys stability and means any bug you hit is yours to patch. Phishing sites: the model card names kokorottsai_com and kokorotts_net as unaffiliated with the model and its author, attaches archive.ph snapshots, and states that any site with kokoro in the root domain implying an affiliation is a red flag. There are only two real entry points, hexgrad/Kokoro-82M on Hugging Face and hexgrad/kokoro on GitHub.

More in Speech & Audio

4
Speech & AudioMediaTopC

Eleven Music v2.5: the closed-source row that turns music generation into a callable pipeline

The music generation model from ElevenLabs, currently at v2.5, released 2026-09-11 with the release post last updated 2026-09-20. We list it as the engineering-side top row for music in the audio domain, which contrasts with rather than duplicates the Suno entry already catalogued here: the value of Suno concentrates inside the product, a Studio multi-track timeline, Custom Models, up to twelve stems and MIDI export, while Eleven Music splits comparable capability into endpoints, namely compose, stream, a structured composition plan, scoring an uploaded video, uploading existing audio, stem separation, finetunes and section level inpainting, plus a marketplace where creators license tracks to each other. Choosing between them therefore needs no audio quality comparison, only one question: is this track something a person sits and adjusts, or something a pipeline requests in volume. The API surface is nine endpoints rather than one generate call, and two details matter for procurement: both compose and stem-separation take a sign_with_c2pa flag that applies to mp3 output so outgoing files can carry content credentials, and output format is tied to subscription tier, with mp3_44100_192 requiring Creator or above and pcm_44100 requiring Pro or above on stem-separation and video-to-music while compose goes up to mp3_48000_320, so the threshold for lossless differs per endpoint. The trap worth memorising is that v2.5 is already the interface default yet the model_id enum of music_v1, music_v2 and music_v2_5 still defaults to music_v1, and output_format=auto resolves per model to mp3_44100_128 on v1 and mp3_48000_192 on v2, so a minimal call that omits model_id silently gets the oldest generation at a lower bitrate; pin both in production code and name mp3_48000_320 when 320kbps is required. The two plan schemas are not interchangeable and the wrong pairing is a hard error: music_v1 takes MusicPrompt while music_v2 and music_v2_5 take CompositionPlan, whose chunks carry a text field with square bracket section names, lyric lines and curly brace inline directions. On the control surface prompt and composition_plan are mutually exclusive, and the request takes music_length_ms from 3000 to 600000, that is three seconds to ten minutes, plus force_instrumental, finetune_id and seed. Rights and commercial use are the most structured part of this line: a multi-year agreement with Universal Music Group was announced alongside v2.5 and the vendor states it is separate from Music 2.5; every track is yours on every plan including Free, Free allows commercial use provided ElevenMusic is credited, lossless downloads are capped at five per day on Free and 400 per month on Pro, tracks built on another artist song through Audio Reference cannot be downloaded, the Marketplace sells licences by usage type with creator earnings starting at 25 percent, and no licence permits distribution to streaming platforms such as Spotify. The official evidence for v2.5 over v2 is a self-run blind test in which v2.5 won the majority of 47,885 paired takes with the widest gap in vocal-led and acoustic-heavy genres. Boundaries: the API is paid-subscription only; vocals are documented for English, Spanish, German and Japanese with no Mandarin, so a Chinese language song is more practical on YuE2 or Suno; seed does not guarantee reproducibility; and it is closed source with no weights, so the self-hostable alternatives here are YuE2 for music and VoiceStudio or Kokoro-82M for speech. Graded C (vendor claim): quality was not recomputed and we ran no blind listening test, but the API contract is verifiable documentation and every clause of it is listed in the body.

10 minAPI 单曲时长上限(music_length_ms 3000-600000)Vendor Claim · 2026-09
ProductionElevenLabsSite
Eleven Music v2.5: the closed-source row that turns music generation into a callable pipeline
Speech & AudioMediaTopA

VibeVoice: a 7.5Hz token rate buys 90 minutes of speech in one pass

The open frontier speech model family from Microsoft, repository microsoft/VibeVoice (MIT, 54,531 stars, verified through the GitHub API). We catalogue it as one asset rather than three because TTS, Realtime and ASR share a single technical core: both the acoustic and the semantic tokenizer are continuous rather than quantised into discrete codebooks, and the frame rate is pushed down to 7.5Hz. That 7.5Hz is where every capability comes from. Ninety minutes of audio at 25Hz is 135,000 tokens and does not fit a 64K context, while at 7.5Hz it is 40,500, so single pass synthesis of 90 minutes and single pass transcription of 60 minutes are two directions of the same fact rather than two separate engineering feats. Five product lines each carry their own ceiling: TTS-1.5B at 90 minutes per run with up to four speakers, Realtime-0.5B at roughly 300ms to first packet, ASR-7B transcribing 60 minutes in one pass while emitting who said what and when with custom hotwords across more than 50 languages, ASR-Streaming emitting text as speech arrives, and ASR-BitNet using heterogeneous quantisation to compress 4.62GB into 1.58GB with RTF under 1 on three or more CPU threads and no GPU. The ASR line is the one worth remembering: conventional long form transcription chains slicing, recognition, diarisation and timestamp alignment, and the global context is lost at the slice boundary, while VibeVoice-ASR fuses recognition, diarisation and timestamps into a single generation, which is the difference between usable and unusable for meeting minutes and support QA. The architecture is next token diffusion, with the LLM owning conversational direction and the diffusion head owning acoustic detail, because turn consistency over a long dialogue is a language problem and only the LLM can hold four speakers distinct across 90 minutes. One fact must be recorded plainly: on 2025-09-05 the team removed the VibeVoice-TTS code from the repository, stating that usage inconsistent with its declared intent had been observed; the weights remain on Hugging Face but the Quick Try section reads Disabled, the official inference scripts are gone and today the TTS path runs through community reproductions, with the model page advising against commercial or real world use without further testing. Every public move since has been on the ASR side, so the commercially usable half is ASR rather than TTS. Its position in tables published by others must also be reported as written: the CosyVoice 3 README lists it at test-zh CER 1.16 with similarity 74.4 and test-en WER 3.04 with similarity 68.9, against CosyVoice3 base at 1.21 and 78.0. Short clip zero shot cloning is not its lane; the three comparisons that matter are how long a single run can be, whether four speakers stay distinct, and whether one hour of meeting audio can be processed in one pass, and on those three the open source field has almost no rival. Boundaries: no official TTS inference path; MIT covers code but not the usage limits on the model page; long form failure is global, since a 90 minute single pass has no retry one segment at a time fallback; the official Risks and Limitations section names deepfakes and requires disclosure of AI involvement. Graded B (confirmed): what is confirmed is the frame rate arithmetic, the published ceilings of each line, the code removal and the licence boundary, not audio quality, which we only relay without recomputing or blind listening.

90 / 60 min单次长音频上限(TTS 4 说话人 / ASR 单遍)Confirmed · 2026-09
ResearchMicrosoftSiteRepo
VibeVoice: a 7.5Hz token rate buys 90 minutes of speech in one pass
Speech & AudioMediaTopC

Fun-CosyVoice3-0.5B: open TTS that turns controllability into an interface

The third generation LLM based speech synthesis system from the FunAudioLLM group at Alibaba, 0.5B parameters, released as Fun-CosyVoice3-0.5B-2512 with base and RL weight sets plus training and inference scripts; the code repository has moved to QwenAudio/CosyVoice (Apache-2.0, 23,794 stars, verified through the GitHub API on 2026-09-29). We list it as the open source top row for speech in the audio domain, and the reason is not that it sounds most human but that it gives each of the four failure modes of production TTS an explicit interface: misread polyphonic characters go through pronunciation correction with Chinese pinyin or English CMU phonemes written straight into the input, wrong readings of numbers and symbols go through the built in text normalisation, collapsed long sentences go through RAS repetition aware sampling, and low latency goes through bidirectional streaming with an official first packet at 150ms, which is the right order of magnitude to sit inside a realtime conversational agent. Three readings matter on the metric sheet. The test-zh speaker similarity of 78.0 is the highest among 0.5B open models and above the human baseline of 75.5, yet still below the closed source Seed-TTS at 79.6, so open source first holds while world first does not. The test-hard column is the real strength, base 6.71 and RL 5.44 being the lowest in the table and better than the closed source 7.59, and hard is exactly long sentences and tongue twisters, so that lead maps to usable real scripts. English needs a discount: base test-en similarity 71.8 sits below the human baseline 73.4 and only the RL tier WER of 1.68 recovers it. Base and RL are two post training weights of one architecture, with all three error rates falling (zh CER down 33 percent, hard CER down 19 percent) and all three similarities slipping slightly, which is a good trade for broadcast and support workloads where one misread character is an incident, while voice fidelity work for a specific IP should audition the RL tier first. The repository also ships GRPO training scripts and a triton plus TensorRT-LLM runtime claimed at four times the speed of HF transformers, so post training is a path others can continue rather than a one off delivery, and the language surface covers nine languages with more than eighteen Chinese dialect accents, a tier the closed APIs still barely offer. Boundaries: the licence is two layered, Apache-2.0 for code while the weight terms live on the model pages and must be checked separately before commercial use; every reading comes from the evaluation of the authors themselves with no third party blind listening table to cross check and we did not recompute; zero shot cloning similarity comes from standard sets while real deployments depend on the noise and style of the reference recording. Graded C (vendor claim).

78.0test-zh 说话人相似度(0.5B 开源档,人类基线 75.5)Vendor Claim · 2025-12
ProductAlibaba Tongyi FunAudioLLM (QwenAudio)SiteRepo
Fun-CosyVoice3-0.5B: open TTS that turns controllability into an interface
Speech & AudioMediaTopC

YuE2-3B: open song generation that exposes the score as an interface

The open song generation model from m-a-p, 3B parameters, weights under CC-BY-NC-4.0, turning lyrics and a style prompt into a complete song with vocals and accompaniment at 48 kHz stereo. We list it as the open-weight top row for music in the audio domain on the strength of a combination that is close to unique among its peers: open weights plus an editable intermediate representation. One AR-NAR Mixture-of-Transformers backbone writes an ABC score (melody and chords) and semantic tokens, flow matching then produces acoustic latents and a VAE decodes them, so the score is an artefact that a person or an agent can read and edit instead of a black box whose only control is another sample. Three cot modes (full, melody, off) map onto composing, covering and direct generation, and the official agentic editing demo runs nine turns across fourteen versions from Mandarin pop to English jazz. Read the benchmark protocol carefully: on 192 WildSongBench prompts the best-of-8 SongBench average of 6.9632 sits above Suno v5 at 6.8721 and Suno v6 at 6.5562, but that is eight candidates with selection against a delivered single candidate; MuLan and AllMusicCaps, the two style-text alignment measures, still favour Suno v5, and PER at 8.44 percent trails Suno v6 Wild at 7.45 and MiniMax Music 3 at 6.27. Cost is the most concrete advantage: a 3.6 minute song in 71 seconds on an RTX 4090 24GB with a peak of 11.18 GiB, and 373 songs per hour on an H800 with vLLM at AR concurrency 32. Limits: non-commercial licence; identity preservation in covers comes almost entirely from a supplied score (CLEWS mAP collapses to 0.006 without one); the benchmark runs on the legacy VAE while the default release is the newer one; no technical report yet and no arena result. Graded C (vendor claim), not recomputed by us.

6.9632WildSongBench SongBench 均分(best-of-8)Vendor Claim · 2026-09
ResearchMultimodal Art Projects (m-a-p)SiteRepo
YuE2-3B: open song generation that exposes the score as an interface