VoiceStudio: the commercial voice stack moved onto local hardware and opened to agents by default
An open source, fully local voice workstation (debpalash/VoiceStudio, AGPL-3.0, Python plus Electron) positioned as the local alternative to ElevenLabs: voice cloning, voice design, video dubbing, dictation, transcription and audiobook creation, with a claimed 646 languages. Created on 2026-04-09 and at 42,831 stars and 5,005 forks on 2026-09-29. We list it as the top self-hosted row in the audio domain, not because any model behind it is stronger, but because the abstraction layer is useful. Open voice capability is scattered across a dozen repositories with their own weight formats, device requirements, cloning interfaces and licences, so assembling a pipeline of clone a voice, transcribe, align, dub and produce an audiobook spends most of its effort on glue. VoiceStudio collects 17 TTS and 11 ASR engines into one registry and reports per engine which devices it runs on (CUDA, MPS, CPU), whether it can clone, and its max_ref_seconds and ref_strategy, on top of a workflow that actually runs. The agent surface matters most for this site: a local API and an MCP server ship as first class product features, the README includes an installation prompt ready to paste into Claude Code, Codex or Cursor, and npx skills add installs it as a skill, which means coding agents can call dubbing and transcription as tools rather than needing a human at the interface. The reference audio policy is equally rigorous, with a 75 second ceiling while how much of a longer clip reaches the model depends on the engine and is reported per engine, an automatic or saved transcript ignored beyond the limit, and a hard [clone_ref_too_long] error when sending more than 20 seconds plus transcript text to the default OmniVoice. One measurement point needs stating plainly: the official docs/benchmarks.md contains a real harness with per stage profiling, refusal to start when memory is insufficient, and a rule that only RTF warm and peak memory on named hardware and versions are accepted with no estimated numbers allowed, but the results column reads No verified rows yet. We therefore attach no performance figure at all and the card carries only the star count we verified ourselves through the GitHub API, which is what earns the A grade. That A asserts the number is real and nothing about being better than ElevenLabs. Four boundaries: AGPL-3.0 is strong copyleft with a network clause, so embedding it in a service offered to outsiders obliges you to release your own server side source, and closed commercial use means either process isolation calling only the local API under your own legal review or assembling the pipeline from Apache-2.0 pieces such as CosyVoice 3 or Kokoro; per-model licences are not settled by the wrapper, since the Breeze-TTS-2 weights behind audio.cpp are research and non-commercial, Supertonic-3 and PocketTTS need extra permissions and GPT-SoVITS requires its own server, so check every model before commercial use; the 646 language figure is the repository own claim and real coverage depends on the engine chosen, from 25 European languages on Parakeet through 50 plus on FunASR to the wider Whisper family, unverified by us language by language; and desktop is Electron only now, with 0.5.3 the last Tauri release, so existing users must migrate. It sits alongside Eleven v4 on the ladder rather than replacing it, one representing the quality ceiling and convenience of a closed commercial stack and the other the data sovereignty and engine substitutability of local self-hosting.