VoiceStudio

VoiceStudio

Fully local ElevenLabs alternative

Description

Dubbing a video into another language, turning a book into an audiobook or generating speech in your own voice at scale gets expensive fast on per-character services like ElevenLabs, and every recording and script has to go to someone else's servers.

VoiceStudio brings the whole toolkit to your own machine: voice cloning, voice design, video dubbing, dictation and transcription all run locally. Give it a clean reference recording to clone a voice and type text to synthesize. Drop in a video and it transcribes, translates and re-voices it in the target language, with original and dub side by side. It covers 646 languages.

It's an open-source desktop app (AGPL-3.0) for Windows, macOS and Linux with over thirty-five thousand GitHub stars. It defaults to the OmniVoice engine and supports others, and it exposes a local API and MCP so AI agents can use it. Only clone voices with the owner's permission.

Features



Voice cloning: reproduce a voice from a reference recording and generate speech from text.

Voice design: create brand-new voices from a description, no reference needed.

Video dubbing: MP4, MKV, MP3 and more, transcribed, translated and re-voiced on a timed track, or paste a video link.

Dictation widget: a floating widget that turns speech into text.

Audiobooks and batches: turn long text and stories into audiobooks and run batch jobs.

Model management: download and switch speech engines in the app to fit your hardware.

Local API and MCP: plug voice features into your own programs or AI agents.