VoiceStudio
1 like
VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.
Features
Properties
- Privacy focused
- Local-First
- AI-Powered
Features
- AI Voice Cloning
- Works Offline
- Dark Mode
- No registration required
- Text to Speech
- Ad-free
- No Tracking
- Voice dictation
- Apple Silicon support
- Local AI
VoiceStudio News & Activities
Highlights All activities
Recent activities
- Kain314 liked VoiceStudio
- POX added VoiceStudio as alternative to Voxtral, VoiceCraft, X to Voice and SherpaTTS
- POX added VoiceStudio
VoiceStudio information
No comments or reviews, maybe you want to be first?
What is VoiceStudio?
Clone voices, dub video, dictate, and produce long-form audio on your own hardware. No account, API key, subscription, or usage meter for the local workflow.
Features:
- Voice Cloning: Zero-shot synthesis from a short reference clip
- Voice Design: Create a voice from age, accent, pitch, style, and delivery instructions
- Video Dubbing: Transcribe, translate, preserve speakers, synthesize, and export video
- Stories and audiobooks: Multi-voice scripts · EPUB/PDF import · chapter rendering · .m4b export
- Dictation Widget: System-wide shortcut, live transcription, optional local-LLM cleanup
- Vocal Isolation: Demucs speech/background separation
- Speaker Diarization: Pyannote and WhisperX speaker assignment
- Batch Queue: Queue large sets of audio and video jobs with per-job progress, or watch a local folder for new videos
- Model Catalogue: Install, remove, select, and route TTS, ASR, and LLM models
- Remote Model Downloads: Install models on enrolled remote workers with live progress
- GPU Auto-Detect: CUDA, MPS, ROCm, and CPU routing with per-engine checks
- AI Watermark: AudioSeal embedding and detection
- MCP Server: Synthesis and transcription tools for MCP clients
- Diagnostics: Self-checks, error journal, logs, and scrubbed support bundles
- Local-first: Core creation stays local; network-backed features are explicit opt-ins
- Extensible: Registry-based TTS, ASR, and plugin interfaces





