Voicebox
Open-source AI voice studio for voice cloning, speech generation, and dictation across seven TTS engines, running locally for free.
Last verified:
What is Voicebox?
Voicebox is an open-source, local AI voice studio that lets you clone voices from minimal audio samples, generate speech across multiple TTS engines, and dictate into any application using AI agents with voices you own. It runs entirely offline on your machine as a free alternative to ElevenLabs and WisprFlow.
Voicebox pricing
Pricing model: Free
Free and open-source
Voicebox pros
- Voice cloning from as little as 3 seconds of audio with natural intonation and emotion
- Completely local/offline with no cloud dependency, subscriptions, or data sharing
- Professional-grade multi-voice timeline editor with audio effects pipeline (reverb, pitch shift, compression, delay)
- Free and open-source with cross-platform support (macOS, Windows, Linux)
Voicebox cons
- Requires local GPU setup and configuration for optimal performance (Metal, CUDA, ROCm, etc.)
- Community-maintained open-source project with limited professional support compared to commercial alternatives
Frequently asked questions about Voicebox
How do I clone a voice?
Upload audio (3+ seconds) via file upload, microphone recording, or system audio capture. Voicebox generates a voice profile automatically.
Does Voicebox require internet or cloud services?
No — everything runs locally on your machine with complete privacy and no cloud dependency or subscriptions.
Can I use dictation system-wide across applications?
Yes — hold Cmd+Alt (macOS) or Ctrl+Alt (Windows) anywhere to dictate; transcripts appear in any app or your clipboard.
What's the maximum length I can generate in one session?
Up to 50,000 characters per generation. Text is auto-split at sentence boundaries and crossfaded seamlessly.