← Blog

How to clone your voice with Voicebox (free, open source, local)

June 29, 2026 · Admin

Clone your voice with Voicebox (free and open source)

Voice cloning used to require expensive software, cloud subscriptions, or technical expertise that put it out of reach for most people. That changed with Voicebox: a free, open-source desktop app that lets you clone your voice and generate speech entirely on your own machine. No cloud uploads, no subscriptions, no accounts required.

What Voicebox is

Voicebox is a local-first AI voice studio created by developer Jamie Pine (the same person behind Spacedrive). It's a free alternative to paid services like ElevenLabs, but everything runs on your hardware. Your voice data, models, and audio never leave your computer.

The app supports seven text-to-speech engines, 23 languages, and can clone a voice from as little as 3 to 5 seconds of audio. It works on macOS, Windows, and Linux, and it's completely free.

Getting started

Go to voicebox.sh and download the installer for your platform. macOS users can grab the Apple Silicon or Intel DMG. Windows users get an MSI installer. Linux users can build from source. The app installs like any standard desktop application with no terminal required.

Once installed, launch Voicebox. The interface is organized into tabs for voices, generation, and a multi-track Stories editor for longer projects.

Recording your reference audio

The key to good voice cloning is quality reference audio. You need a clean recording of the voice you want to clone.

Record in a quiet room since background noise degrades cloning quality. Use a decent microphone, though even a smartphone mic works if the environment is quiet. Speak naturally and don't exaggerate or whisper. Aim for 10 to 30 seconds since longer samples produce better results, though Voicebox can work with just a few seconds. Avoid music or other voices in the reference recording.

You can record directly in Voicebox or import an existing audio file (WAV, MP3, or M4A). See the audio guide for tips on recording clean narration if you're new to this.

Creating a voice profile

In the Voices tab, click "Import Voice" or "Create Voice." Upload your reference audio, give the profile a name, and select a TTS engine. Voicebox recommends Qwen3-TTS for the best cloning quality: it's the primary engine and achieves near-perfect results from short samples.

Once the profile is created, you'll see it in your voice list. You can create multiple profiles for different voices and switch between them instantly.

Generating speech

Select your cloned voice, type or paste your text into the text field, and hit generate. The audio renders locally, and speed depends on your hardware, but modern machines handle it in seconds. You can adjust pitch, reverb, delay, and other effects in the post-processing panel.

For longer scripts, Voicebox auto-chunks the text with crossfade for seamless output. The Stories editor lets you layer multiple voices on a timeline for podcasts, dialogues, or multi-character narration.

Why local matters

Privacy is the headline advantage. Cloud-based cloning services like ElevenLabs send your voice data to external servers. Voicebox keeps everything local, making it viable for sensitive use cases (voiceovers, accessibility tools, personal projects) without surrendering control of your audio identity.

At zero cost, Voicebox has made professional-grade voice cloning accessible to anyone with a decent computer. Quality varies by TTS engine and hardware, but it's a significant step toward democratizing speech synthesis.

Limitations to know about

Voicebox requires a machine with enough RAM and compute to run AI models locally. Older or low-spec hardware will experience slow generation. The quality ceiling is also below what cloud services like ElevenLabs offer on premium tiers, though the gap is narrowing.

For short-form content where you just need fast, browser-based TTS without your own voice, Voicertool or NoteGPT are faster to use with no setup. See the no-signup TTS tools comparison for a breakdown of when each makes sense.

From cloned voice to finished video

Once you have your generated audio from Voicebox, the next step is turning it into a publishable video. The fastest path is uploading the MP3 to wavreel: it transcribes the speech, matches stock footage from Pexels to each scene, adds animated captions, and renders a finished vertical or horizontal MP4.

This combination (Voicebox for voice generation, wavreel for video assembly) covers the full faceless content pipeline without any paid subscriptions or editing skills required.

For the full workflow from script to published video, see the faceless channel workflow guide.

Related tools and guides


Upload your Voicebox audio and get a finished video with wavreel →