AIToolScan

Voicebox

Voicebox Overview

Voicebox is an open-source AI voice studio that runs entirely on your local machine — a free, local alternative to ElevenLabs and WisprFlow for voice cloning, dictation, and agent voice output. With 1M+ downloads, it bundles seven TTS engines and lets you clone, dictate, and create in voices you own.

  • Voice Cloning: Clones any voice from as little as a 3-second audio sample with natural intonation and emotion — from an uploaded clip, microphone, or system audio (YouTube, podcasts).
  • 7 TTS Engines: Bundles Qwen3-TTS, Chatterbox, Chatterbox Turbo, LuxTTS, Qwen CustomVoice, TADA, and Kokoro out of the box, plus Whisper STT (99 languages) and Qwen3 for transcript refinement.
  • Stories Editor: Create multi-voice narratives with a timeline-based editor — arrange tracks, trim clips, and mix conversations between characters.
  • Voices with Personality: Give any voice profile a free-form personality, then Rewrite your text or Compose new lines entirely in their voice, in full character.
  • Audio Effects Pipeline: Apply pitch shift, reverb, delay, compression and more, then save presets and set defaults per voice profile.
  • Dictate Anywhere: Hold a global shortcut to dictate into any app on macOS or Windows, powered by Whisper STT with refined transcripts that clean up ums and self-corrections.
  • MCP Integration: Any MCP-aware agent (Claude Code, Cursor, Cline) speaks through your cloned voices with one tool call (voicebox.speak).
  • Local REST API: Every engine becomes a localhost endpoint for speech generation, profiles and health — no API keys, no rate limits, no per-character fees, unlimited 50,000-character generations.