Skip to content
←Back to Open Source

OPEN SOURCE DEEP DIVE

SpeechToTextSubtitleTranslationAIDubbing

SmartSub: Open-Source One-Stop Desktop Tool for Subtitle Generation, Translation, AI Dubbing & Voice Cloning

SmartSub is an open-source desktop tool (5.5k Stars, MIT license) that packs the entire audio/video subtitle and dubbing pipeline — speech-to-text, translation, proofreading, TTS dubbing, burn-in — into one app. Supports 8 transcription engines (whisper.cpp/faster-whisper/FunASR/Qwen3-ASR/FireRedASR/Parakeet etc.), 20 translation services, local TTS with zero-shot voice cloning. Built-in AI assistant (multimodal visual diagnostics + Agentic tool-calling) and 111 MCP tools + CLI commands for Cursor/Claude Code/Codex integration. Full pipeline runs completely free, cross-platform Windows/macOS/Linux with full GPU acceleration.

buxuku/SmartSub5.5kTypeScriptMIT6 min read

Overview

SmartSub is an open-source desktop tool for audio/video subtitles and dubbing, packing the entire pipeline — speech-to-text → subtitle translation → proofreading → TTS dubbing → burn-in composition — into one application. 5.5k Stars, MIT license, cross-platform on Windows / macOS / Linux, with GPU acceleration for NVIDIA CUDA, AMD/Intel Vulkan, and Apple Core ML/Metal. The entire pipeline can run completely free: local model transcription, built-in free translation sources, local TTS dubbing (including voice cloning), and local ffmpeg burn-in — no API Key required.

Core Capabilities

Online Video Download

Paste Bilibili, YouTube, or other platform links to download directly. Dual-engine (yt-dlp supporting 1800+ sites, lux for Chinese platforms) auto-matches by platform. Can simultaneously download official platform subtitles (including auto-generated), auto-paired in the task wizard — if official subtitles exist, no transcription needed. Supports importing site cookies to unlock HD quality and member content, stored locally only.

Subtitle Generation (Transcription)

Supports 8 transcription engines switchable per task:

  • whisper.cpp (built-in): Default engine, supports ggml quantized models and GPU acceleration, ships with the app
  • faster-whisper: Based on CTranslate2, faster, models downloaded from HuggingFace on demand
  • FunASR: SenseVoice (Chinese/English/Japanese/Korean/Cantonese) and Paraformer-zh, excellent Chinese performance, runs via built-in sherpa-onnx native library
  • Qwen3-ASR: Tongyi Qianwen speech recognition (qwen3-asr-0.6b / 1.7b)
  • FireRedASR: FireRedASR-AED large, excellent Chinese-English performance
  • NVIDIA Parakeet: English TDT v2, multilingual TDT v3 (25 European languages), Japanese 0.6B CTC
  • Local Whisper CLI: Calls your own installed whisper-compatible command
  • Cloud transcription: 9 online service providers (OpenAI-compatible, ElevenLabs Scribe, Deepgram, Volcano Engine Doubao, Tencent Cloud, Alibaba Cloud, iFlytek, Gladia, Xiaomi MiMo), no GPU needed, multi-provider multi-instance

AI subtitle refinement (optional): LLM semantic segmentation + batch correction — segmentation reorganizes by semantics while word-level timestamps remain precise; correction fixes homophones, removes filler words, normalizes punctuation. Provider defaults to AI translation config (local Ollama at zero cost), auto-falls back to rule-based segmentation on failure.

Subtitle Translation

20 translation services: built-in free translation (Bing/Google free endpoints with auto-fallback and rate limiting), Baidu, Alibaba Cloud, Tencent, iFlytek, Volcano Engine, Doubao, Niutrans, DeepLX, Azure, Google, plus LLM services including Ollama (local), DeepSeek, Gemini, Tongyi Qianwen, SiliconFlow, Azure OpenAI, and DeerAPI. Compatible with any OpenAI-style API. Outputs pure translation or bilingual "original + translation" subtitles. Each AI service supports custom request parameters (String/Float/Boolean/Array/Object/Integer types) with import/export.

Subtitle Proofreading

Built-in proofreading workbench for line-by-line video-referenced checking with precise positioning. Undo/redo support, recoverable single-line deletion. AI one-click polish, plus on-demand AI assistant for conversational editing and Q&A.

TTS Dubbing & Voice Cloning

Independent dubbing workbench: one subtitle file + optional video, line-by-line speech synthesis with automatic timeline alignment.

Local engines (offline, free):

  • Kokoro multilingual v1.1: Chinese/English, 103 voices
  • VITS Chinese AIShell3: Chinese, 174 speaker voices
  • ZipVoice voice cloning: zero-shot cloning, one reference audio + corresponding text

Cloud services: Edge TTS (free), OpenAI-compatible endpoints, Azure Speech (700+ voices), Volcano Engine Doubao (Voice Replication 2.0), ElevenLabs (instant cloning), Xiaomi MiMo.

Timeline alignment mechanism: pre-controls speech rate before synthesis, verifies actual duration after synthesis (local engines re-synthesize free, cloud uses atempo speed change), borrows time from adjacent silent gaps when insufficient; lines exceeding 1.5x speed threshold enter manual processing queue. Reference audio auto-quality-checked (duration, SNR, clipping, volume) when creating cloned voices.

Video Composition (Subtitle Burn-in)

Hard subtitles (permanently burned into frame) and soft subtitles (lossless stream-copy packaging with switchable subtitle tracks). Font, size, color, outline, shadow, 9-grid positioning with multiple preset styles, WYSIWYG real-time preview.

AI Creative Assistant (Copilot)

SmartSub deeply integrates modern AI collaboration and Agentic automation:

  • Instant summon: Hotkey ⌘J (macOS) / Ctrl+J (Windows/Linux) opens the right drawer panel, session history auto-persisted locally
  • Deep workspace context awareness: Intelligently senses current UI state, auto-associates video playback position, selected subtitle line, neighboring context, task list status, and recent error logs
  • Agentic tool-calling: Fully connected to 111 underlying automation tools — can directly execute tasks from natural language instructions like "convert current video to bilingual Chinese-English subtitles and compose"
  • Subtitle refinement & safe editing: Say "make this line more colloquial" to edit, real-time rendering in editor with undo/redo stack, strict version protection — only writes to file on explicit "save"
  • Multimodal visual screenshot diagnostics: Click camera icon to capture workspace, combined with vision multimodal LLM to analyze error popups and control states, giving troubleshooting in seconds
  • Multi-type file & media chat: Drag audio/video, subtitles, reference scripts, and images into chat; audio/video files stay local, only paths passed to underlying tools
  • Flexible LLM integration: Supports DeepSeek, Tongyi Qianwen, Gemini, SiliconFlow, DeerAPI, and any OpenAI-compatible endpoint

MCP Protocol & CLI Automation

  • 111 standard MCP tools with corresponding CLI commands: Comprehensive coverage of video download, transcription, translation, proofreading, dubbing, packaging, format conversion, audio extraction, and model/service configuration
  • Zero-config Node.js runtime: Reuses the app's built-in Electron/Node runtime — clients and external AI need no Node.js installation
  • One-click AI client integration: Cursor (one-click import), Codex (TOML config), Claude Code (one-click registration), and generic client JSON config
  • Headless daemon: Auto-starts a windowless background daemon on first AI tool or CLI call, seamlessly shares task state, config, and model resources with the desktop client, auto-reclaims when idle
  • Powerful CLI: Unified smartsub entry point, supports pipeline.run orchestration, JSON/stdin pipe parameters, task status polling and dedup

Full Pipeline Free Plan

StageFree OptionNotes
Video downloadyt-dlp / lux open-source enginesOne-click install in-app
Transcriptionwhisper.cpp / faster-whisper / FunASR / Qwen3-ASR / FireRedASR / ParakeetDownload model once, use offline
TranslationBuilt-in free (Bing/Google), Ollama local LLM, DeepLXZero-config, out of the box
TTS dubbingLocal Kokoro / VITS / ZipVoice cloning; Edge TTS free tierLocal models, no usage limits
Burn-inBuilt-in ffmpegLocal composition

Technical Architecture

Built on Nextron (Next.js + Electron), primarily JavaScript/TypeScript. Transcription engines run via whisper.cpp native addon and sherpa-onnx native library — no Python environment needed. GPU acceleration supports NVIDIA CUDA (11.8.0/12.2.0/12.4.0/13.0.2), AMD/Intel Vulkan, Apple Core ML/Metal, with in-app accelerator package download and automatic CPU fallback on failure.

Installation

Windows/Linux x64 and macOS (Apple Silicon/Intel) installers available. macOS recommended via Homebrew:

brew tap buxuku/tap
brew install --cask smartsub
brew upgrade --cask smartsub

Three steps to start: install and download a speech model → select task, drag in files or paste links → begin processing.

Acknowledgments & Community

Built on whisper.cpp (local transcription engine), sherpa-onnx (runtime for FunASR/Qwen3-ASR/FireRedASR/Parakeet and local TTS), and FFmpeg (audio/video processing and subtitle burn-in). QQ community group: 655348339. MIT License.