
Description
Dubbing a video into another language, turning a book into an audiobook or generating speech in your own voice at scale gets expensive fast on per-character services like ElevenLabs, and every recording and script has to go to someone else's servers.
VoiceStudio brings the whole toolkit to your own machine: voice cloning, voice design, video dubbing, dictation and transcription all run locally. Give it a clean reference recording to clone a voice and type text to synthesize. Drop in a video and it transcribes, translates and re-voices it in the target language, with original and dub side by side. It covers 646 languages.
It's an open-source desktop app (AGPL-3.0) for Windows, macOS and Linux with over thirty-five thousand GitHub stars. It defaults to the OmniVoice engine and supports others, and it exposes a local API and MCP so AI agents can use it. Only clone voices with the owner's permission.
Voice cloning: reproduce a voice from a reference recording and generate speech from text.
Voice design: create brand-new voices from a description, no reference needed.
Video dubbing: MP4, MKV, MP3 and more, transcribed, translated and re-voiced on a timed track, or paste a video link.
Dictation widget: a floating widget that turns speech into text.
Audiobooks and batches: turn long text and stories into audiobooks and run batch jobs.
Model management: download and switch speech engines in the app to fit your hardware.
Local API and MCP: plug voice features into your own programs or AI agents.
VoiceStudio brings the whole toolkit to your own machine: voice cloning, voice design, video dubbing, dictation and transcription all run locally. Give it a clean reference recording to clone a voice and type text to synthesize. Drop in a video and it transcribes, translates and re-voices it in the target language, with original and dub side by side. It covers 646 languages.
It's an open-source desktop app (AGPL-3.0) for Windows, macOS and Linux with over thirty-five thousand GitHub stars. It defaults to the OmniVoice engine and supports others, and it exposes a local API and MCP so AI agents can use it. Only clone voices with the owner's permission.
Features
Voice cloning: reproduce a voice from a reference recording and generate speech from text.
Voice design: create brand-new voices from a description, no reference needed.
Video dubbing: MP4, MKV, MP3 and more, transcribed, translated and re-voiced on a timed track, or paste a video link.
Dictation widget: a floating widget that turns speech into text.
Audiobooks and batches: turn long text and stories into audiobooks and run batch jobs.
Model management: download and switch speech engines in the app to fit your hardware.
Local API and MCP: plug voice features into your own programs or AI agents.

