VoiceStudio: the commercial voice stack moved onto local hardware and opened to agents by default
voicestudio
An open source, fully local voice workstation (debpalash/VoiceStudio, AGPL-3.0, Python plus Electron) positioned as the local alternative to ElevenLabs: voice cloning, voice design, video dubbing, dictation, transcription and audiobook creation, with a claimed 646 languages. Created on 2026-04-09 and at 42,831 stars and 5,005 forks on 2026-09-29. We list it as the top self-hosted row in the audio domain, not because any model behind it is stronger, but because the abstraction layer is useful. Open voice capability is scattered across a dozen repositories with their own weight formats, device requirements, cloning interfaces and licences, so assembling a pipeline of clone a voice, transcribe, align, dub and produce an audiobook spends most of its effort on glue. VoiceStudio collects 17 TTS and 11 ASR engines into one registry and reports per engine which devices it runs on (CUDA, MPS, CPU), whether it can clone, and its max_ref_seconds and ref_strategy, on top of a workflow that actually runs. The agent surface matters most for this site: a local API and an MCP server ship as first class product features, the README includes an installation prompt ready to paste into Claude Code, Codex or Cursor, and npx skills add installs it as a skill, which means coding agents can call dubbing and transcription as tools rather than needing a human at the interface. The reference audio policy is equally rigorous, with a 75 second ceiling while how much of a longer clip reaches the model depends on the engine and is reported per engine, an automatic or saved transcript ignored beyond the limit, and a hard [clone_ref_too_long] error when sending more than 20 seconds plus transcript text to the default OmniVoice. One measurement point needs stating plainly: the official docs/benchmarks.md contains a real harness with per stage profiling, refusal to start when memory is insufficient, and a rule that only RTF warm and peak memory on named hardware and versions are accepted with no estimated numbers allowed, but the results column reads No verified rows yet. We therefore attach no performance figure at all and the card carries only the star count we verified ourselves through the GitHub API, which is what earns the A grade. That A asserts the number is real and nothing about being better than ElevenLabs. Four boundaries: AGPL-3.0 is strong copyleft with a network clause, so embedding it in a service offered to outsiders obliges you to release your own server side source, and closed commercial use means either process isolation calling only the local API under your own legal review or assembling the pipeline from Apache-2.0 pieces such as CosyVoice 3 or Kokoro; per-model licences are not settled by the wrapper, since the Breeze-TTS-2 weights behind audio.cpp are research and non-commercial, Supertonic-3 and PocketTTS need extra permissions and GPT-SoVITS requires its own server, so check every model before commercial use; the 646 language figure is the repository own claim and real coverage depends on the engine chosen, from 25 European languages on Parakeet through 50 plus on FunASR to the wider Whisper family, unverified by us language by language; and desktop is Electron only now, with 0.5.3 the last Tauri release, so existing users must migrate. It sits alongside Eleven v4 on the ladder rather than replacing it, one representing the quality ceiling and convenience of a closed commercial stack and the other the data sovereignty and engine substitutability of local self-hosting.
- CONFIDENCE
- Confirmed
- Two or more independent sources, or reproduced by our harness
- KEY METRIC
- GitHub Stars(2026-09-29,本站经 GitHub API 核)
- Confirmed · 2026-09
- MATURITY
- Product
- research → demo → product → production
Our takeWe grade it A (verifiable), with a precise statement of what the A covers. The only number on the card is 42,831 GitHub stars as of 2026-09-29, pulled by us directly from the GitHub API, and because we verified that reading ourselves it is recorded as confirmed under contract section 13.5. It carries no performance conclusion whatsoever. The official
docs/benchmarks.mdlists No verified rows yet in its results column, with only the harness description and contribution guidance above it. We therefore attach no RTF, memory or quality figure to this entry, not one. That is the largest measurement difference between this row and the others in the same domain, and it is part of why the entry deserves a place here: a benchmark harness that genuinely exists but is left empty is more honest than an estimated score sheet. The inverse reading also holds, and matters for selection: anyone wanting performance data to justify choosing VoiceStudio has nothing citable, from us or from the project.What earns it the top of the self-hosted row in the audio domain is not a stronger model but a useful abstraction layer. Open voice capability is scattered across a dozen repositories, each with its own weight format, device requirements, cloning interface and licence, so assembling a pipeline of clone a voice, transcribe, align, dub, produce an audiobook spends most of its effort on glue. VoiceStudio collects 17 TTS and 11 ASR engines into one registry and reports per engine which devices it runs on (CUDA, MPS, CPU), whether it can clone, and its
max_ref_secondsandref_strategy, on top of a workflow that actually runs. The agent surface is the part that matters most for this site: a local API and an MCP server ship as first class product features, the README includes an installation prompt ready to paste into Claude Code, Codex or Cursor, andnpx skills addinstalls it as a skill, which means the coding agents catalogued in our harness domain can call dubbing and transcription as tools rather than needing a human at the interface. The reference audio policy shows the same unusual rigour: a 75 second ceiling, but how much of a longer clip actually reaches the model depends on the engine, stated in the UI and reported per engine through the API; beyond the limit an automatic or saved transcript is ignored, and sending more than 20 seconds plus transcript text to the default OmniVoice fails with[clone_ref_too_long]. Exposing how much an engine truncates as a machine readable field instead of leaving users to discover it by trial is the dividing line between a demo and a product.Four hard boundaries govern selection. AGPL-3.0 is strong copyleft with a network clause, so embedding it in a service offered to outsiders obliges you to release your own server side source under AGPL. Using it inside a closed commercial product means either process isolation that only calls the local API, which needs your own legal review, or assembling the same pipeline from Apache-2.0 pieces such as CosyVoice 3 or Kokoro, which is why this site keeps assembled products and individual models on separate ladder rows. Per-model licences are not settled by the AGPL wrapper: the Breeze-TTS-2 weights behind audio.cpp are research and non-commercial, Supertonic-3 and PocketTTS need extra permissions, GPT-SoVITS requires running its own server, and every model must be checked before commercial use. The 646 language figure is the repository own claim, and real coverage depends on the engine chosen, from 25 European languages on Parakeet through 50 plus on FunASR to the wider Whisper family; we have not verified it language by language. Desktop is Electron only now, since 0.5.3 was the last Tauri release and the old shell plus legacy UI entry points were removed, so existing users must migrate.
A closing note on how to read the star count. 42.8k stars in five months, a repository created on 2026-04-09, and only 41 open issues. Stars measure attention rather than quality, and a low issue count in a fast growing repository can mean responsive maintenance just as easily as a user base that has not yet settled into filing reports. Our confirmed grade asserts only that the number is real, not that this product is better than ElevenLabs. It sits alongside Eleven v4 on the ladder rather than replacing it: one represents the quality ceiling and convenience of a closed commercial stack, the other the data sovereignty and engine substitutability of local self-hosting, and the line between them is whether your data may leave the machine and whether you need to hold the voice weights yourself.
The problem it solves: moving the whole commercial voice workflow onto local hardware and exposing it to agents
VoiceStudio (debpalash/VoiceStudio, AGPL-3.0, Python plus Electron) describes itself as the open source, fully local alternative to ElevenLabs: voice cloning, voice design, video dubbing, dictation, transcription and audiobook creation, with a claimed 646 languages. The repository was created on 2026-04-09 and by 2026-09-29 had reached 42,831 stars and 5,005 forks with pushes still landing most days. We list it as the top self-hosted row in the audio domain, not because any single model behind it is stronger, but because it collects open voice capabilities scattered across a dozen repositories into one product with an actual workflow, and opens that workflow to agents by default.
Three things hold at once and separate it from a closed service: it runs locally (workflows execute on your hardware, remote workers are optional and usage analytics requires consent), there is no subscription, and engines are swappable. The closed stack wins on convenience and consistency; VoiceStudio wins on data never leaving the machine, replaceable models and a one-off hardware cost.
The engine catalogue: 17 TTS and 11 ASR backends laid out by device and cloning support
This is the most intelligence-dense part of the project. It does not train its own model so much as wire in what the open ecosystem already has, then document what each engine runs on, whether it clones and how to enable it:
| Class | Engine | Runs on | Cloning |
|---|---|---|---|
| TTS, default and mainline | VoiceStudio / OmniVoice (default, from k2-fsa) | CUDA · MPS · CPU | Yes |
| VoxCPM2 | CUDA · MPS · CPU | Yes, plus voice design | |
| CosyVoice 3 | CUDA · CPU | Yes | |
| IndexTTS 2.5 (one-click sidecar install) | CUDA · CPU | Yes, plus emotion | |
| TTS, lightweight and edge | KittenTTS (8 preset voices) | CPU only | No |
| Sherpa-ONNX | CUDA · CPU | No | |
| Supertonic-3 (7 preset voices) | CPU only | No | |
| MLX-Audio (Kokoro, CSM, Dia and others) | Apple Silicon | Model dependent | |
| TTS, bring your own clone or extra licence | GPT-SoVITS (through its own API server) | External server | Yes |
| MOSS-TTS-v1.5 (8B), dots.tts (2B), MOSS-TTS-Nano, Confucius4-TTS | CUDA · CPU | Yes | |
| PocketTTS (Kyutai), audio.cpp (Breeze-TTS-2, weights research and non-commercial) | CPU | Yes, plus voice design | |
| ASR | WhisperX (default, word timestamps plus diarisation, built for dubbing) | CUDA · CPU | - |
| Faster-Whisper (including an isolated unattended batch mode), MLX Whisper, PyTorch Whisper for ROCm hosts | By platform | - | |
| Parakeet TDT (NeMo and MLX, 25 European languages), Moonshine (edge, no timestamps), FunASR SenseVoice (50+ languages, inline diarisation), Sherpa-ONNX live dictation, OpenAI-compatible endpoints | By engine | - |
Speaker diarisation is not an engine registry of its own: the dubbing pipeline uses pyannote (a gated Hugging Face model) and FunASR can diarise inline with its cam++ speaker model. The feature list also covers voice design, video dubbing with timed speech, a floating dictation widget, vocal isolation, a batch queue, an AI watermark and GPU auto-detection across CUDA, ROCm, MPS and CPU, pinnable in settings or through OMNIVOICE_DEVICE.
Reference audio: a 75 second ceiling, but how much reaches the model is per engine
This is unusually rigorous for a project of its kind. Voice Clone accepts clips up to 75 seconds and recommends 5 to 15 seconds of clean speech, and how much of a longer clip reaches the model depends on the engine. The screen says which applies and GET /engines reports it per engine as max_ref_seconds and ref_strategy. The default OmniVoice uses up to 20 seconds with a best_window strategy that selects the 15 second passage containing the most speech and transcribes it; VoxCPM2 keeps the first 30 seconds after edge-silence trimming (head); the rest are marked unverified and pass the clip through (null).
One trap follows from that design. A transcript of the whole clip cannot match the passage the engine actually cuts out, so automatic and saved transcripts are ignored when a clip exceeds the limit, and sending a transcript with an OmniVoice clip longer than 20 seconds is rejected with [clone_ref_too_long]. The fix is to trim audio and transcript to the same passage or omit the transcript. The limits are identical when an engine runs in its own environment.
Agent native: a local API, an MCP server, and an install prompt you can paste into a coding agent
Agent support here is deeper than having an endpoint. A local API and an MCP server ship as first class product features, the README hands you a prompt to paste into Claude Code, Codex or Cursor, and docs/install/agent.md covers hardware detection, reusing existing data, asking before downloading models and running a test generation. Agents with skill support can run npx skills add debpalash/VoiceStudio and choose between a voicestudio skill for audio workflows and a voicestudio-maintainer skill for repository work. In practice this means the coding agents catalogued under agent harness on this site can call dubbing, transcription and audiobook production as tools instead of driving a GUI.
Limits: AGPL, an empty benchmark table, per-model licences, and the Tauri to Electron migration
- AGPL-3.0 is strong copyleft with a network clause. Embedding it in a service you offer to others obliges you to release your own server side source under AGPL. Using it inside a closed commercial product means process isolation with only local API calls (and your own legal review), or assembling Apache-2.0 components such as CosyVoice 3 and Kokoro instead.
- The official benchmark table is currently empty.
docs/benchmarks.mdcontains a real harness (scripts/bench_pipeline.py, profiling each pipeline stage one at a time, refusing to start a stage without enough free RAM and unloading models between stages) and states that it collects only warm RTF and peak VRAM produced by that harness on named hardware at a named version, with nothing estimated. The results section reads No verified rows yet and contains only contribution instructions. This site therefore publishes no performance figure for it and uses verifiable GitHub stars as the card metric. That is simultaneously the honest part (no estimates passed off as measurements) and the current evidence gap. - Each model carries its own licence, so AGPL on the app does not make everything usable. The Breeze-TTS-2 weights behind audio.cpp are research and non-commercial, Supertonic-3 and PocketTTS need additional licences, and GPT-SoVITS requires running its own server. Every model needs a licence check before commercial use, and voice cloning is permitted only with the consent of the person cloned.
- The 646 language figure is a repository claim. Real coverage depends on the selected engine, and the engines differ widely (Parakeet covers 25 European languages, FunASR 50+, the Whisper family more). We have not verified it language by language.
- Electron is now the only desktop app. Version 0.5.3 was the final Tauri release, the retired Tauri shell and legacy UI entry points have been removed, and existing Tauri users must install Electron separately using the migration guide. The install script needs curl and a SHA-256 tool; building from main needs Git, Node.js 22+, Bun, Rust and Cargo plus platform build tools.
- The repository is young and grew very fast. 42.8k stars in five months, against only 41 open issues in a repository created on 2026-04-09. Stars measure attention rather than quality, so selection should be based on whether the engine catalogue and the local API plus MCP abstraction are useful to you, not on trending position.