Latest commit

History

1,893 Commits

Folders and files

NameName
Last commit message
Last commit date
VoiceStudio ranking on TrendshiftVoiceStudio logo

VoiceStudio

Previously OmniVoice-Studio

Clone voices, dub video, dictate, and produce long-form audio on your own hardware.

16 TTS engines · 11 ASR engines · 646-language catalogue · macOS, Windows, Linux, and Docker

No account, API key, subscription, or usage meter for the local workflow.

Install · Features · Compare · Requirements · Engines · Architecture · API · Docs · 简体中文

GitHub starsTotal downloadsLatest releaseAGPL-3.0 licenseDiscord community

Download VoiceStudio

Switching TTS engines from the VoiceStudio status bar

Warning

Active beta. Use the latest release for stable work. main contains the newest fixes and may change between releases. Report problems through GitHub Issues.

At a glance

VoiceStudio
WorkflowsVoice cloning and design, video dubbing, dictation, stories, audiobooks, batch generation
Language catalogue646 TTS languages; actual coverage and quality depend on the selected engine
Engines16 TTS · 11 ASR · switch in Model Catalogue or with Ctrl/Cmd+E
PlatformsmacOS 13.3+ on Apple Silicon · Windows 10/11 x64 · Linux x86_64 with glibc 2.39+
ComputeCUDA · Apple Silicon MPS/MLX · ROCm on Linux · CPU · optional remote workers
InterfacesDesktop app · local REST/SSE/WebSocket API · OpenAI-compatible audio API · MCP Server
StorageVoices, projects, settings, and outputs stay on the machine by default
LicenseAGPL-3.0 application; downloaded models keep their upstream terms

Install

Download a package from the latest release, then follow the platform guide.

PlatformPackageGuide
macOS 13.3+Apple Silicon DMGInstall on macOS
Windows 10/11x64 MSI; choose the current-user build when listed to install without admin accessInstall on Windows
LinuxAppImage, x86_64 with glibc 2.39+Install on Linux
DockerCUDA, ROCm, CPU, and worker-only GPU profilesRun with Docker

First launch creates a managed Python environment and downloads the default model. Later launches reuse both.

Note

On macOS, first launch needs a one-time right-click, then Open approval. Intel Macs cannot run the local Python backend; use a remote backend instead.

First voice

  1. Launch VoiceStudio and open Voice Cloning.
  2. Add a clean voice sample. Three seconds works; 5 to 15 seconds usually gives a better prompt.
  3. Enter text, choose a language, then select Generate.

Run from source

Install the development prerequisites, then:

git clone https://github.com/debpalash/VoiceStudio.git
cd VoiceStudio
bun install
bun run desktop

Use bun run dev for the browser UI. See Contributing for services, tests, and platform packages.

If setup fails

Features

AreaIncluded
Voice CloningZero-shot synthesis from a short reference clip
Voice DesignCreate a voice from age, accent, pitch, style, and delivery instructions
Video DubbingTranscribe, translate, preserve speakers, synthesize, and export video
Stories and audiobooksMulti-voice scripts · EPUB/PDF import · chapter rendering · .m4b export
Dictation WidgetSystem-wide shortcut, live transcription, optional local-LLM cleanup
Vocal IsolationDemucs speech/background separation
Speaker DiarizationPyannote and WhisperX speaker assignment
Batch QueueQueue large sets of audio and video jobs with per-job progress
Model CatalogueInstall, remove, select, and route TTS, ASR, and LLM models
Remote Model DownloadsInstall models on enrolled remote workers with live progress
GPU Auto-DetectCUDA, MPS, ROCm, and CPU routing with per-engine checks
AI WatermarkAudioSeal embedding and detection
MCP ServerSynthesis and transcription tools for MCP clients
DiagnosticsSelf-checks, error journal, logs, and scrubbed support bundles
Local-firstCore creation stays local; network-backed features are explicit opt-ins
ExtensibleRegistry-based TTS, ASR, and plugin interfaces
VoiceStudio Model CatalogueSaving a gallery voice as a local profile
Model Catalogue: engine, device, and install stateGallery: save a shared voice as a local profile

Comparison

VoiceStudio trades managed cloud compute for local control. This is the practical difference:

VoiceStudioTypical hosted voice service
Best fitPrivate, offline, self-hosted, or high-volume workFast setup without local model management
Data pathLocal by default; remote features are opt-inAudio and text are processed by the provider
Cost modelFree software; you supply the hardwareSubscription, credits, or metered API use
SetupInstall the app and model weightsCreate an account and use the web app or API
PerformanceDepends on your engine and hardwareProvider manages compute and scaling
Offline useYes, after required models are installedUsually requires a network connection
CustomizationSource, engines, models, API, and routing are openLimited to provider options
MaintenanceYou manage updates, disk, and computeProvider manages infrastructure

Requirements

Requirements vary by engine. These values cover the default local workflow.

MinimumRecommended
OSWindows 10 x64 · macOS 13.3 Apple Silicon · Linux x86_64 with glibc 2.39+Current supported OS release
RAM8 GB16 GB+
Disk10 GB free20 GB+ SSD
GPUOptional; CPU mode is supportedNVIDIA CUDA or Apple Silicon
VRAM4 GB when using a GPU8 GB+; large optional engines need more
Python from source3.11+3.11 or 3.12

ROCm is Linux-only and opt-in. Windows AMD/Ryzen AI uses CPU. Systems with limited VRAM offload work to CPU when required. See performance, benchmarks, and engine disk usage.

Engines

Engine support is capability-specific. Check cloning, language, platform, memory, and license before choosing one. Full setup guides: docs/engines.

Text to speech

EngineLanguagesCloneInstructLinuxmacOS ARMWindowsLicense
VoiceStudio (default, powered by k2-fsa/OmniVoice)600+YesYesCUDA/CPUMPSCUDA/CPUAGPL-3.0 app · Apache-2.0 code, CC-BY-NC weights³
CosyVoice 39 + 18 dialectsYesYesCUDA/CPUCPUCUDA/CPUApache-2.0
GPT-SoVITS5YesNoCUDA/CPUNoCUDA/CPUMIT
VoxCPM230YesYesCUDA/CPUMPSCUDA/CPUApache-2.0
MOSS-TTS-Nano20YesNoCUDA/CPUCPUCUDA/CPUApache-2.0
KittenTTSEnglishNoNoCPUCPUCPUMIT
MLX-AudioModel-dependentVariesVariesNoMLXNoVaries
Sherpa-ONNX20+NoNoCUDA/CPUCPUCUDA/CPUApache-2.0
IndexTTS 2.5ZH · EN · JA · ES · ARYesNoCUDA/CPUCPUCUDA/CPUBilibili model license¹
OmniVoice GGUF600+YesYesCUDA/CPUMPS/CPUCUDA/CPUAGPL-3.0 app · review the derivative model terms³
OmniVoice (subprocess)600+YesYesCUDA/CPUMPSCUDA/CPUAGPL-3.0 app · Apache-2.0 code, CC-BY-NC weights³
PocketTTSEN · FR · DE · PT · IT · ESYesNoCPUCPUCPUCC-BY-4.0, gated²
Supertonic 331NoNoCPUCPUCPUOpenRAIL-M
MOSS-TTS-v1.531YesNoCUDA/CPUCPUCUDA/CPUApache-2.0
dots.tts24YesNoCUDA/CPUCPUNoApache-2.0
Confucius4-TTS14YesNoCUDA/CPUCPUCUDA/CPUApache-2.0

⚡ Installed or registered on demand.

¹ IndexTTS 2.5 requires a separate written Bilibili license above 100 million monthly active users or RMB 1 billion annual revenue. Review the model license.

² PocketTTS shows its gated-access and CC-BY-4.0 terms before first use.

³ The OmniVoice snapshot also includes an audio tokenizer under separate Boson Higgs Audio 2 and Meta Llama community terms. VoiceStudio's application license does not replace model or tokenizer terms.

Clone-less engines cannot preserve a reference speaker in dubbing or pinned-voice batch jobs. VoiceStudio rejects those jobs instead of silently changing engines. Heavy engines have separate memory and platform limits; check their engine guide first.

Speech to text

EngineIDLanguagesBest fit
WhisperX (default)whisperx~100Dubbing, subtitles, word-level timing
Faster-Whisperfaster-whisper~100General cross-platform transcription
Faster-Whisper (isolated)faster-whisper-isolated~100Crash-isolated batch transcription
MLX Whispermlx-whisper~100Apple Silicon
PyTorch Whisperpytorch-whisper~100CUDA, MPS, and CPU fallback
Parakeet TDTnemo-parakeetEnglish + 25 EUFast CPU/CUDA transcription
Parakeet TDT v3 (MLX)parakeet-mlx25 EUApple Silicon dictation and word timestamps
MoonshinemoonshineEnglishLow-power, low-latency ONNX
FunASRfunasr50+VAD and inline diarization
sherpa-onnx (live dictation)sherpa-onnx-asrModel-dependentStreaming CPU dictation
OpenAI-compatible⚠️ configured serveropenai-compat-asrServer-dependentLocal gigastt/Qwen3-ASR or a remote endpoint; audio goes only to that server

WhisperX and Faster-Whisper retry with int8 when efficient float16 is unavailable. Pin ASR_COMPUTE_TYPE=int8 or float32 only if automatic selection still fails.

Architecture

Tauri v2 desktop shell (Rust)
│ IPC
React + Vite UI
│ HTTP · SSE · WebSocket on localhost:3900
FastAPI backend
├── TTS / ASR engine registries
├── dubbing / audio / long-form pipelines
├── OpenAI-compatible API and MCP server
└── SQLite + Alembic → omnivoice_data/
LayerPathResponsibility
Desktop shellfrontend/src-tauri/Window lifecycle, tray, shortcuts, updater, sidecar bootstrap
Frontendfrontend/src/React UI, Zustand state, API and event clients, i18n
APIbackend/api/REST routes, schemas, auth boundaries, streaming
Core servicesbackend/services/Generation, dubbing, audio processing, persistence
Enginesbackend/engines/Isolated and optional engine adapters
Worker systembackend/worker/Authenticated remote compute and job transport
Dataomnivoice_data/Projects, voices, settings, logs, and SQLite state
Deliveryscripts/, deploy/, .github/workflows/Development, packaging, containers, releases, CI

Network boundary

  • The desktop talks to a loopback-only backend on localhost:3900.
  • Loopback API calls need no server key. Remote access requires a share PIN or API key.
  • Remote workers and OpenAI-compatible ASR are opt-in. Loopback ASR may use HTTP and keeps audio on the machine; non-loopback endpoints require HTTPS, and redirects are not followed.
  • Analytics is off until consent. If enabled, it sends allowlisted, content-free usage metadata. It never sends text, audio, file names, or projects.

Local speech platform and OpenAI-compatible API

Point an OpenAI-compatible audio client at the local backend:

- base_url="https://api.openai.com/v1"+ base_url="http://localhost:3900/v1"
EndpointPurpose
POST /v1/audio/speechTTS to mp3, opus, aac, flac, wav, or pcm; select a profile with voice and an engine with model
POST /v1/audio/transcriptionsSTT to json, text, verbose_json, srt, or vtt
WS /v1/audio/transcriptions/streamLive PCM/WebM transcription with partial, utterance, and session-final events
GET /.well-known/voicestudio-speechDiscover HTTP, WebSocket, MCP, and native dictation-control transports
GET /v1/audio/voicesList local voice profiles and engines
fromopenaiimportOpenAIclient=OpenAI(base_url="http://localhost:3900/v1", api_key="local")
withclient.audio.speech.with_streaming_response.create(
model="tts-1",
voice="<profile-id>",
input="Made on my own hardware.",
response_format="wav",
) asresponse:
response.stream_to_file("speech.wav")

The bundled Rust control sidecar lets Herdr, coding agents, VS Code, desktop apps, and TUIs trigger the system-wide dictation flow or reuse its native text insertion. See the speech platform guide. The full API reference is in Settings → OpenAPI Reference. For LAN, Tailscale, or proxy access, read API authentication before exposing the backend.

Agent skills

Install the VoiceStudio skills for Claude Code, Codex, Cursor, and other skills.sh-compatible agents:

npx skills add debpalash/VoiceStudio
  • omnivoice: synthesize speech and transcribe audio through local VoiceStudio.
  • oss-maintainer: the repository's open-source maintenance workflow.

Google Colab

Open in Colab

The notebook runs the app and web UI on a Colab GPU. Colab is remote compute, so uploaded audio and project data do not remain local to your machine.

Documentation

NeedRead
InstallmacOS · Windows · Linux · Docker
Fix setupTroubleshooting · model downloads · Hugging Face token
Choose an engineEngine guides · benchmarks · expressive speech
Tune hardwarePerformance · remote workers
Build integrationsSpeech platform · Private production API · API auth · MCP · examples
Build VoiceStudioContributing · engine acceptance
Track changesChangelog · roadmap · latest release
Remove everythingUninstall guide

FAQ

Does it work on Apple Silicon and Intel Macs?

Apple Silicon is supported with MPS and MLX options. Intel Macs cannot run the local backend because current PyTorch wheels are unavailable; they can connect to a remote backend. See macOS installation.

How much VRAM do I need?

A GPU is optional. Use 4 GB VRAM as the minimum for accelerated work and 8 GB+ for the default multi-stage workflow. Large optional engines can require 12 to 16 GB or more. Check the benchmarks and engine guide.

Why does a longer reference clip not always improve the clone?

Cloning is zero-shot: the clip is a prompt, not training data. Use 5 to 15 seconds of one speaker, close to the microphone, without music, noise, or reverb. Match the tone and pace you want in the output. For training, see data preparation and training.

Can I use generated audio commercially?

VoiceStudio's application license does not restrict generated audio, but it does not grant rights under a model's separate terms. The default OmniVoice repository labels its pretrained weights CC-BY-NC and includes a tokenizer under separate community terms. Review the selected model terms before commercial use.

Does VoiceStudio collect data?

Not unless you opt in. Analytics is off by default and skipping consent keeps it off. When enabled, the app sends allowlisted, content-free usage metadata. Text, audio, file names, voices, and projects are excluded. Change this at Settings → Privacy.

How do I remove VoiceStudio and its data?

Use scripts/uninstall.sh on macOS/Linux or scripts\uninstall.ps1 on Windows. Both show a dry run before deletion. See the uninstall guide for every path.

Community and contributing

Support development

VoiceStudio is free and has no paid tier. Donations fund development and infrastructure.

Ko-fi · PayPal · Sponsorship details

License

VoiceStudio is licensed under AGPL-3.0. You may run it, modify it, and use it internally. The application license itself does not restrict selling generated audio, but downloaded model and tokenizer terms may. If you modify VoiceStudio and provide that modified version as a network service, AGPL requires you to offer the corresponding source under the same license. A commercial license for VoiceStudio-owned code is available for proprietary embedding; it does not relicense third-party models. Contact VoiceStudio@palash.dev. See LICENSE-NOTICE.md for the plain-language scope.

Optional engines and downloaded models retain their own licenses. The bundled omnivoice/ Python code is Apache-2.0 upstream; the default downloaded weights and audio tokenizer use separate terms.

Acknowledgments

VoiceStudio builds on OmniVoice, WhisperX, Demucs, Pyannote, CTranslate2, AudioSeal, Tauri, Supertonic, Sherpa-ONNX, GPT-SoVITS, and PocketTTS.

About

VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

14.0k stars

Watchers

67 watching

Forks

Releases

Sponsor this project

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

1,893 Commits

Folders and files

NameName
Last commit message
Last commit date
VoiceStudio ranking on TrendshiftVoiceStudio logo

VoiceStudio

Previously OmniVoice-Studio

Clone voices, dub video, dictate, and produce long-form audio on your own hardware.

16 TTS engines · 11 ASR engines · 646-language catalogue · macOS, Windows, Linux, and Docker

No account, API key, subscription, or usage meter for the local workflow.

Install · Features · Compare · Requirements · Engines · Architecture · API · Docs · 简体中文

GitHub starsTotal downloadsLatest releaseAGPL-3.0 licenseDiscord community

Download VoiceStudio

Switching TTS engines from the VoiceStudio status bar

Warning

Active beta. Use the latest release for stable work. main contains the newest fixes and may change between releases. Report problems through GitHub Issues.

At a glance

VoiceStudio
WorkflowsVoice cloning and design, video dubbing, dictation, stories, audiobooks, batch generation
Language catalogue646 TTS languages; actual coverage and quality depend on the selected engine
Engines16 TTS · 11 ASR · switch in Model Catalogue or with Ctrl/Cmd+E
PlatformsmacOS 13.3+ on Apple Silicon · Windows 10/11 x64 · Linux x86_64 with glibc 2.39+
ComputeCUDA · Apple Silicon MPS/MLX · ROCm on Linux · CPU · optional remote workers
InterfacesDesktop app · local REST/SSE/WebSocket API · OpenAI-compatible audio API · MCP Server
StorageVoices, projects, settings, and outputs stay on the machine by default
LicenseAGPL-3.0 application; downloaded models keep their upstream terms

Install

Download a package from the latest release, then follow the platform guide.

PlatformPackageGuide
macOS 13.3+Apple Silicon DMGInstall on macOS
Windows 10/11x64 MSI; choose the current-user build when listed to install without admin accessInstall on Windows
LinuxAppImage, x86_64 with glibc 2.39+Install on Linux
DockerCUDA, ROCm, CPU, and worker-only GPU profilesRun with Docker

First launch creates a managed Python environment and downloads the default model. Later launches reuse both.

Note

On macOS, first launch needs a one-time right-click, then Open approval. Intel Macs cannot run the local Python backend; use a remote backend instead.

First voice

  1. Launch VoiceStudio and open Voice Cloning.
  2. Add a clean voice sample. Three seconds works; 5 to 15 seconds usually gives a better prompt.
  3. Enter text, choose a language, then select Generate.

Run from source

Install the development prerequisites, then:

git clone https://github.com/debpalash/VoiceStudio.git
cd VoiceStudio
bun install
bun run desktop

Use bun run dev for the browser UI. See Contributing for services, tests, and platform packages.

If setup fails

Features

AreaIncluded
Voice CloningZero-shot synthesis from a short reference clip
Voice DesignCreate a voice from age, accent, pitch, style, and delivery instructions
Video DubbingTranscribe, translate, preserve speakers, synthesize, and export video
Stories and audiobooksMulti-voice scripts · EPUB/PDF import · chapter rendering · .m4b export
Dictation WidgetSystem-wide shortcut, live transcription, optional local-LLM cleanup
Vocal IsolationDemucs speech/background separation
Speaker DiarizationPyannote and WhisperX speaker assignment
Batch QueueQueue large sets of audio and video jobs with per-job progress
Model CatalogueInstall, remove, select, and route TTS, ASR, and LLM models
Remote Model DownloadsInstall models on enrolled remote workers with live progress
GPU Auto-DetectCUDA, MPS, ROCm, and CPU routing with per-engine checks
AI WatermarkAudioSeal embedding and detection
MCP ServerSynthesis and transcription tools for MCP clients
DiagnosticsSelf-checks, error journal, logs, and scrubbed support bundles
Local-firstCore creation stays local; network-backed features are explicit opt-ins
ExtensibleRegistry-based TTS, ASR, and plugin interfaces
VoiceStudio Model CatalogueSaving a gallery voice as a local profile
Model Catalogue: engine, device, and install stateGallery: save a shared voice as a local profile

Comparison

VoiceStudio trades managed cloud compute for local control. This is the practical difference:

VoiceStudioTypical hosted voice service
Best fitPrivate, offline, self-hosted, or high-volume workFast setup without local model management
Data pathLocal by default; remote features are opt-inAudio and text are processed by the provider
Cost modelFree software; you supply the hardwareSubscription, credits, or metered API use
SetupInstall the app and model weightsCreate an account and use the web app or API
PerformanceDepends on your engine and hardwareProvider manages compute and scaling
Offline useYes, after required models are installedUsually requires a network connection
CustomizationSource, engines, models, API, and routing are openLimited to provider options
MaintenanceYou manage updates, disk, and computeProvider manages infrastructure

Requirements

Requirements vary by engine. These values cover the default local workflow.

MinimumRecommended
OSWindows 10 x64 · macOS 13.3 Apple Silicon · Linux x86_64 with glibc 2.39+Current supported OS release
RAM8 GB16 GB+
Disk10 GB free20 GB+ SSD
GPUOptional; CPU mode is supportedNVIDIA CUDA or Apple Silicon
VRAM4 GB when using a GPU8 GB+; large optional engines need more
Python from source3.11+3.11 or 3.12

ROCm is Linux-only and opt-in. Windows AMD/Ryzen AI uses CPU. Systems with limited VRAM offload work to CPU when required. See performance, benchmarks, and engine disk usage.

Engines

Engine support is capability-specific. Check cloning, language, platform, memory, and license before choosing one. Full setup guides: docs/engines.

Text to speech

EngineLanguagesCloneInstructLinuxmacOS ARMWindowsLicense
VoiceStudio (default, powered by k2-fsa/OmniVoice)600+YesYesCUDA/CPUMPSCUDA/CPUAGPL-3.0 app · Apache-2.0 code, CC-BY-NC weights³
CosyVoice 39 + 18 dialectsYesYesCUDA/CPUCPUCUDA/CPUApache-2.0
GPT-SoVITS5YesNoCUDA/CPUNoCUDA/CPUMIT
VoxCPM230YesYesCUDA/CPUMPSCUDA/CPUApache-2.0
MOSS-TTS-Nano20YesNoCUDA/CPUCPUCUDA/CPUApache-2.0
KittenTTSEnglishNoNoCPUCPUCPUMIT
MLX-AudioModel-dependentVariesVariesNoMLXNoVaries
Sherpa-ONNX20+NoNoCUDA/CPUCPUCUDA/CPUApache-2.0
IndexTTS 2.5ZH · EN · JA · ES · ARYesNoCUDA/CPUCPUCUDA/CPUBilibili model license¹
OmniVoice GGUF600+YesYesCUDA/CPUMPS/CPUCUDA/CPUAGPL-3.0 app · review the derivative model terms³
OmniVoice (subprocess)600+YesYesCUDA/CPUMPSCUDA/CPUAGPL-3.0 app · Apache-2.0 code, CC-BY-NC weights³
PocketTTSEN · FR · DE · PT · IT · ESYesNoCPUCPUCPUCC-BY-4.0, gated²
Supertonic 331NoNoCPUCPUCPUOpenRAIL-M
MOSS-TTS-v1.531YesNoCUDA/CPUCPUCUDA/CPUApache-2.0
dots.tts24YesNoCUDA/CPUCPUNoApache-2.0
Confucius4-TTS14YesNoCUDA/CPUCPUCUDA/CPUApache-2.0

⚡ Installed or registered on demand.

¹ IndexTTS 2.5 requires a separate written Bilibili license above 100 million monthly active users or RMB 1 billion annual revenue. Review the model license.

² PocketTTS shows its gated-access and CC-BY-4.0 terms before first use.

³ The OmniVoice snapshot also includes an audio tokenizer under separate Boson Higgs Audio 2 and Meta Llama community terms. VoiceStudio's application license does not replace model or tokenizer terms.

Clone-less engines cannot preserve a reference speaker in dubbing or pinned-voice batch jobs. VoiceStudio rejects those jobs instead of silently changing engines. Heavy engines have separate memory and platform limits; check their engine guide first.

Speech to text

EngineIDLanguagesBest fit
WhisperX (default)whisperx~100Dubbing, subtitles, word-level timing
Faster-Whisperfaster-whisper~100General cross-platform transcription
Faster-Whisper (isolated)faster-whisper-isolated~100Crash-isolated batch transcription
MLX Whispermlx-whisper~100Apple Silicon
PyTorch Whisperpytorch-whisper~100CUDA, MPS, and CPU fallback
Parakeet TDTnemo-parakeetEnglish + 25 EUFast CPU/CUDA transcription
Parakeet TDT v3 (MLX)parakeet-mlx25 EUApple Silicon dictation and word timestamps
MoonshinemoonshineEnglishLow-power, low-latency ONNX
FunASRfunasr50+VAD and inline diarization
sherpa-onnx (live dictation)sherpa-onnx-asrModel-dependentStreaming CPU dictation
OpenAI-compatible⚠️ configured serveropenai-compat-asrServer-dependentLocal gigastt/Qwen3-ASR or a remote endpoint; audio goes only to that server

WhisperX and Faster-Whisper retry with int8 when efficient float16 is unavailable. Pin ASR_COMPUTE_TYPE=int8 or float32 only if automatic selection still fails.

Architecture

Tauri v2 desktop shell (Rust)
│ IPC
React + Vite UI
│ HTTP · SSE · WebSocket on localhost:3900
FastAPI backend
├── TTS / ASR engine registries
├── dubbing / audio / long-form pipelines
├── OpenAI-compatible API and MCP server
└── SQLite + Alembic → omnivoice_data/
LayerPathResponsibility
Desktop shellfrontend/src-tauri/Window lifecycle, tray, shortcuts, updater, sidecar bootstrap
Frontendfrontend/src/React UI, Zustand state, API and event clients, i18n
APIbackend/api/REST routes, schemas, auth boundaries, streaming
Core servicesbackend/services/Generation, dubbing, audio processing, persistence
Enginesbackend/engines/Isolated and optional engine adapters
Worker systembackend/worker/Authenticated remote compute and job transport
Dataomnivoice_data/Projects, voices, settings, logs, and SQLite state
Deliveryscripts/, deploy/, .github/workflows/Development, packaging, containers, releases, CI

Network boundary

  • The desktop talks to a loopback-only backend on localhost:3900.
  • Loopback API calls need no server key. Remote access requires a share PIN or API key.
  • Remote workers and OpenAI-compatible ASR are opt-in. Loopback ASR may use HTTP and keeps audio on the machine; non-loopback endpoints require HTTPS, and redirects are not followed.
  • Analytics is off until consent. If enabled, it sends allowlisted, content-free usage metadata. It never sends text, audio, file names, or projects.

Local speech platform and OpenAI-compatible API

Point an OpenAI-compatible audio client at the local backend:

- base_url="https://api.openai.com/v1"+ base_url="http://localhost:3900/v1"
EndpointPurpose
POST /v1/audio/speechTTS to mp3, opus, aac, flac, wav, or pcm; select a profile with voice and an engine with model
POST /v1/audio/transcriptionsSTT to json, text, verbose_json, srt, or vtt
WS /v1/audio/transcriptions/streamLive PCM/WebM transcription with partial, utterance, and session-final events
GET /.well-known/voicestudio-speechDiscover HTTP, WebSocket, MCP, and native dictation-control transports
GET /v1/audio/voicesList local voice profiles and engines
fromopenaiimportOpenAIclient=OpenAI(base_url="http://localhost:3900/v1", api_key="local")
withclient.audio.speech.with_streaming_response.create(
model="tts-1",
voice="<profile-id>",
input="Made on my own hardware.",
response_format="wav",
) asresponse:
response.stream_to_file("speech.wav")

The bundled Rust control sidecar lets Herdr, coding agents, VS Code, desktop apps, and TUIs trigger the system-wide dictation flow or reuse its native text insertion. See the speech platform guide. The full API reference is in Settings → OpenAPI Reference. For LAN, Tailscale, or proxy access, read API authentication before exposing the backend.

Agent skills

Install the VoiceStudio skills for Claude Code, Codex, Cursor, and other skills.sh-compatible agents:

npx skills add debpalash/VoiceStudio
  • omnivoice: synthesize speech and transcribe audio through local VoiceStudio.
  • oss-maintainer: the repository's open-source maintenance workflow.

Google Colab

Open in Colab

The notebook runs the app and web UI on a Colab GPU. Colab is remote compute, so uploaded audio and project data do not remain local to your machine.

Documentation

NeedRead
InstallmacOS · Windows · Linux · Docker
Fix setupTroubleshooting · model downloads · Hugging Face token
Choose an engineEngine guides · benchmarks · expressive speech
Tune hardwarePerformance · remote workers
Build integrationsSpeech platform · Private production API · API auth · MCP · examples
Build VoiceStudioContributing · engine acceptance
Track changesChangelog · roadmap · latest release
Remove everythingUninstall guide

FAQ

Does it work on Apple Silicon and Intel Macs?

Apple Silicon is supported with MPS and MLX options. Intel Macs cannot run the local backend because current PyTorch wheels are unavailable; they can connect to a remote backend. See macOS installation.

How much VRAM do I need?

A GPU is optional. Use 4 GB VRAM as the minimum for accelerated work and 8 GB+ for the default multi-stage workflow. Large optional engines can require 12 to 16 GB or more. Check the benchmarks and engine guide.

Why does a longer reference clip not always improve the clone?

Cloning is zero-shot: the clip is a prompt, not training data. Use 5 to 15 seconds of one speaker, close to the microphone, without music, noise, or reverb. Match the tone and pace you want in the output. For training, see data preparation and training.

Can I use generated audio commercially?

VoiceStudio's application license does not restrict generated audio, but it does not grant rights under a model's separate terms. The default OmniVoice repository labels its pretrained weights CC-BY-NC and includes a tokenizer under separate community terms. Review the selected model terms before commercial use.

Does VoiceStudio collect data?

Not unless you opt in. Analytics is off by default and skipping consent keeps it off. When enabled, the app sends allowlisted, content-free usage metadata. Text, audio, file names, voices, and projects are excluded. Change this at Settings → Privacy.

How do I remove VoiceStudio and its data?

Use scripts/uninstall.sh on macOS/Linux or scripts\uninstall.ps1 on Windows. Both show a dry run before deletion. See the uninstall guide for every path.

Community and contributing

Support development

VoiceStudio is free and has no paid tier. Donations fund development and infrastructure.

Ko-fi · PayPal · Sponsorship details

License

VoiceStudio is licensed under AGPL-3.0. You may run it, modify it, and use it internally. The application license itself does not restrict selling generated audio, but downloaded model and tokenizer terms may. If you modify VoiceStudio and provide that modified version as a network service, AGPL requires you to offer the corresponding source under the same license. A commercial license for VoiceStudio-owned code is available for proprietary embedding; it does not relicense third-party models. Contact VoiceStudio@palash.dev. See LICENSE-NOTICE.md for the plain-language scope.

Optional engines and downloaded models retain their own licenses. The bundled omnivoice/ Python code is Apache-2.0 upstream; the default downloaded weights and audio tokenizer use separate terms.

Acknowledgments

VoiceStudio builds on OmniVoice, WhisperX, Demucs, Pyannote, CTranslate2, AudioSeal, Tauri, Supertonic, Sherpa-ONNX, GPT-SoVITS, and PocketTTS.

About

VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

14.0k stars

Watchers

67 watching

Forks

Releases

Sponsor this project

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

1,893 Commits

Folders and files

NameName
Last commit message
Last commit date
VoiceStudio ranking on TrendshiftVoiceStudio logo

VoiceStudio

Previously OmniVoice-Studio

Clone voices, dub video, dictate, and produce long-form audio on your own hardware.

16 TTS engines · 11 ASR engines · 646-language catalogue · macOS, Windows, Linux, and Docker

No account, API key, subscription, or usage meter for the local workflow.

Install · Features · Compare · Requirements · Engines · Architecture · API · Docs · 简体中文

GitHub starsTotal downloadsLatest releaseAGPL-3.0 licenseDiscord community

Download VoiceStudio

Switching TTS engines from the VoiceStudio status bar

Warning

Active beta. Use the latest release for stable work. main contains the newest fixes and may change between releases. Report problems through GitHub Issues.

At a glance

VoiceStudio
WorkflowsVoice cloning and design, video dubbing, dictation, stories, audiobooks, batch generation
Language catalogue646 TTS languages; actual coverage and quality depend on the selected engine
Engines16 TTS · 11 ASR · switch in Model Catalogue or with Ctrl/Cmd+E
PlatformsmacOS 13.3+ on Apple Silicon · Windows 10/11 x64 · Linux x86_64 with glibc 2.39+
ComputeCUDA · Apple Silicon MPS/MLX · ROCm on Linux · CPU · optional remote workers
InterfacesDesktop app · local REST/SSE/WebSocket API · OpenAI-compatible audio API · MCP Server
StorageVoices, projects, settings, and outputs stay on the machine by default
LicenseAGPL-3.0 application; downloaded models keep their upstream terms

Install

Download a package from the latest release, then follow the platform guide.

PlatformPackageGuide
macOS 13.3+Apple Silicon DMGInstall on macOS
Windows 10/11x64 MSI; choose the current-user build when listed to install without admin accessInstall on Windows
LinuxAppImage, x86_64 with glibc 2.39+Install on Linux
DockerCUDA, ROCm, CPU, and worker-only GPU profilesRun with Docker

First launch creates a managed Python environment and downloads the default model. Later launches reuse both.

Note

On macOS, first launch needs a one-time right-click, then Open approval. Intel Macs cannot run the local Python backend; use a remote backend instead.

First voice

  1. Launch VoiceStudio and open Voice Cloning.
  2. Add a clean voice sample. Three seconds works; 5 to 15 seconds usually gives a better prompt.
  3. Enter text, choose a language, then select Generate.

Run from source

Install the development prerequisites, then:

git clone https://github.com/debpalash/VoiceStudio.git
cd VoiceStudio
bun install
bun run desktop

Use bun run dev for the browser UI. See Contributing for services, tests, and platform packages.

If setup fails

Features

AreaIncluded
Voice CloningZero-shot synthesis from a short reference clip
Voice DesignCreate a voice from age, accent, pitch, style, and delivery instructions
Video DubbingTranscribe, translate, preserve speakers, synthesize, and export video
Stories and audiobooksMulti-voice scripts · EPUB/PDF import · chapter rendering · .m4b export
Dictation WidgetSystem-wide shortcut, live transcription, optional local-LLM cleanup
Vocal IsolationDemucs speech/background separation
Speaker DiarizationPyannote and WhisperX speaker assignment
Batch QueueQueue large sets of audio and video jobs with per-job progress
Model CatalogueInstall, remove, select, and route TTS, ASR, and LLM models
Remote Model DownloadsInstall models on enrolled remote workers with live progress
GPU Auto-DetectCUDA, MPS, ROCm, and CPU routing with per-engine checks
AI WatermarkAudioSeal embedding and detection
MCP ServerSynthesis and transcription tools for MCP clients
DiagnosticsSelf-checks, error journal, logs, and scrubbed support bundles
Local-firstCore creation stays local; network-backed features are explicit opt-ins
ExtensibleRegistry-based TTS, ASR, and plugin interfaces
VoiceStudio Model CatalogueSaving a gallery voice as a local profile
Model Catalogue: engine, device, and install stateGallery: save a shared voice as a local profile

Comparison

VoiceStudio trades managed cloud compute for local control. This is the practical difference:

VoiceStudioTypical hosted voice service
Best fitPrivate, offline, self-hosted, or high-volume workFast setup without local model management
Data pathLocal by default; remote features are opt-inAudio and text are processed by the provider
Cost modelFree software; you supply the hardwareSubscription, credits, or metered API use
SetupInstall the app and model weightsCreate an account and use the web app or API
PerformanceDepends on your engine and hardwareProvider manages compute and scaling
Offline useYes, after required models are installedUsually requires a network connection
CustomizationSource, engines, models, API, and routing are openLimited to provider options
MaintenanceYou manage updates, disk, and computeProvider manages infrastructure

Requirements

Requirements vary by engine. These values cover the default local workflow.

MinimumRecommended
OSWindows 10 x64 · macOS 13.3 Apple Silicon · Linux x86_64 with glibc 2.39+Current supported OS release
RAM8 GB16 GB+
Disk10 GB free20 GB+ SSD
GPUOptional; CPU mode is supportedNVIDIA CUDA or Apple Silicon
VRAM4 GB when using a GPU8 GB+; large optional engines need more
Python from source3.11+3.11 or 3.12

ROCm is Linux-only and opt-in. Windows AMD/Ryzen AI uses CPU. Systems with limited VRAM offload work to CPU when required. See performance, benchmarks, and engine disk usage.

Engines

Engine support is capability-specific. Check cloning, language, platform, memory, and license before choosing one. Full setup guides: docs/engines.

Text to speech

EngineLanguagesCloneInstructLinuxmacOS ARMWindowsLicense
VoiceStudio (default, powered by k2-fsa/OmniVoice)600+YesYesCUDA/CPUMPSCUDA/CPUAGPL-3.0 app · Apache-2.0 code, CC-BY-NC weights³
CosyVoice 39 + 18 dialectsYesYesCUDA/CPUCPUCUDA/CPUApache-2.0
GPT-SoVITS5YesNoCUDA/CPUNoCUDA/CPUMIT
VoxCPM230YesYesCUDA/CPUMPSCUDA/CPUApache-2.0
MOSS-TTS-Nano20YesNoCUDA/CPUCPUCUDA/CPUApache-2.0
KittenTTSEnglishNoNoCPUCPUCPUMIT
MLX-AudioModel-dependentVariesVariesNoMLXNoVaries
Sherpa-ONNX20+NoNoCUDA/CPUCPUCUDA/CPUApache-2.0
IndexTTS 2.5ZH · EN · JA · ES · ARYesNoCUDA/CPUCPUCUDA/CPUBilibili model license¹
OmniVoice GGUF600+YesYesCUDA/CPUMPS/CPUCUDA/CPUAGPL-3.0 app · review the derivative model terms³
OmniVoice (subprocess)600+YesYesCUDA/CPUMPSCUDA/CPUAGPL-3.0 app · Apache-2.0 code, CC-BY-NC weights³
PocketTTSEN · FR · DE · PT · IT · ESYesNoCPUCPUCPUCC-BY-4.0, gated²
Supertonic 331NoNoCPUCPUCPUOpenRAIL-M
MOSS-TTS-v1.531YesNoCUDA/CPUCPUCUDA/CPUApache-2.0
dots.tts24YesNoCUDA/CPUCPUNoApache-2.0
Confucius4-TTS14YesNoCUDA/CPUCPUCUDA/CPUApache-2.0

⚡ Installed or registered on demand.

¹ IndexTTS 2.5 requires a separate written Bilibili license above 100 million monthly active users or RMB 1 billion annual revenue. Review the model license.

² PocketTTS shows its gated-access and CC-BY-4.0 terms before first use.

³ The OmniVoice snapshot also includes an audio tokenizer under separate Boson Higgs Audio 2 and Meta Llama community terms. VoiceStudio's application license does not replace model or tokenizer terms.

Clone-less engines cannot preserve a reference speaker in dubbing or pinned-voice batch jobs. VoiceStudio rejects those jobs instead of silently changing engines. Heavy engines have separate memory and platform limits; check their engine guide first.

Speech to text

EngineIDLanguagesBest fit
WhisperX (default)whisperx~100Dubbing, subtitles, word-level timing
Faster-Whisperfaster-whisper~100General cross-platform transcription
Faster-Whisper (isolated)faster-whisper-isolated~100Crash-isolated batch transcription
MLX Whispermlx-whisper~100Apple Silicon
PyTorch Whisperpytorch-whisper~100CUDA, MPS, and CPU fallback
Parakeet TDTnemo-parakeetEnglish + 25 EUFast CPU/CUDA transcription
Parakeet TDT v3 (MLX)parakeet-mlx25 EUApple Silicon dictation and word timestamps
MoonshinemoonshineEnglishLow-power, low-latency ONNX
FunASRfunasr50+VAD and inline diarization
sherpa-onnx (live dictation)sherpa-onnx-asrModel-dependentStreaming CPU dictation
OpenAI-compatible⚠️ configured serveropenai-compat-asrServer-dependentLocal gigastt/Qwen3-ASR or a remote endpoint; audio goes only to that server

WhisperX and Faster-Whisper retry with int8 when efficient float16 is unavailable. Pin ASR_COMPUTE_TYPE=int8 or float32 only if automatic selection still fails.

Architecture

Tauri v2 desktop shell (Rust)
│ IPC
React + Vite UI
│ HTTP · SSE · WebSocket on localhost:3900
FastAPI backend
├── TTS / ASR engine registries
├── dubbing / audio / long-form pipelines
├── OpenAI-compatible API and MCP server
└── SQLite + Alembic → omnivoice_data/
LayerPathResponsibility
Desktop shellfrontend/src-tauri/Window lifecycle, tray, shortcuts, updater, sidecar bootstrap
Frontendfrontend/src/React UI, Zustand state, API and event clients, i18n
APIbackend/api/REST routes, schemas, auth boundaries, streaming
Core servicesbackend/services/Generation, dubbing, audio processing, persistence
Enginesbackend/engines/Isolated and optional engine adapters
Worker systembackend/worker/Authenticated remote compute and job transport
Dataomnivoice_data/Projects, voices, settings, logs, and SQLite state
Deliveryscripts/, deploy/, .github/workflows/Development, packaging, containers, releases, CI

Network boundary

  • The desktop talks to a loopback-only backend on localhost:3900.
  • Loopback API calls need no server key. Remote access requires a share PIN or API key.
  • Remote workers and OpenAI-compatible ASR are opt-in. Loopback ASR may use HTTP and keeps audio on the machine; non-loopback endpoints require HTTPS, and redirects are not followed.
  • Analytics is off until consent. If enabled, it sends allowlisted, content-free usage metadata. It never sends text, audio, file names, or projects.

Local speech platform and OpenAI-compatible API

Point an OpenAI-compatible audio client at the local backend:

- base_url="https://api.openai.com/v1"+ base_url="http://localhost:3900/v1"
EndpointPurpose
POST /v1/audio/speechTTS to mp3, opus, aac, flac, wav, or pcm; select a profile with voice and an engine with model
POST /v1/audio/transcriptionsSTT to json, text, verbose_json, srt, or vtt
WS /v1/audio/transcriptions/streamLive PCM/WebM transcription with partial, utterance, and session-final events
GET /.well-known/voicestudio-speechDiscover HTTP, WebSocket, MCP, and native dictation-control transports
GET /v1/audio/voicesList local voice profiles and engines
fromopenaiimportOpenAIclient=OpenAI(base_url="http://localhost:3900/v1", api_key="local")
withclient.audio.speech.with_streaming_response.create(
model="tts-1",
voice="<profile-id>",
input="Made on my own hardware.",
response_format="wav",
) asresponse:
response.stream_to_file("speech.wav")

The bundled Rust control sidecar lets Herdr, coding agents, VS Code, desktop apps, and TUIs trigger the system-wide dictation flow or reuse its native text insertion. See the speech platform guide. The full API reference is in Settings → OpenAPI Reference. For LAN, Tailscale, or proxy access, read API authentication before exposing the backend.

Agent skills

Install the VoiceStudio skills for Claude Code, Codex, Cursor, and other skills.sh-compatible agents:

npx skills add debpalash/VoiceStudio
  • omnivoice: synthesize speech and transcribe audio through local VoiceStudio.
  • oss-maintainer: the repository's open-source maintenance workflow.

Google Colab

Open in Colab

The notebook runs the app and web UI on a Colab GPU. Colab is remote compute, so uploaded audio and project data do not remain local to your machine.

Documentation

NeedRead
InstallmacOS · Windows · Linux · Docker
Fix setupTroubleshooting · model downloads · Hugging Face token
Choose an engineEngine guides · benchmarks · expressive speech
Tune hardwarePerformance · remote workers
Build integrationsSpeech platform · Private production API · API auth · MCP · examples
Build VoiceStudioContributing · engine acceptance
Track changesChangelog · roadmap · latest release
Remove everythingUninstall guide

FAQ

Does it work on Apple Silicon and Intel Macs?

Apple Silicon is supported with MPS and MLX options. Intel Macs cannot run the local backend because current PyTorch wheels are unavailable; they can connect to a remote backend. See macOS installation.

How much VRAM do I need?

A GPU is optional. Use 4 GB VRAM as the minimum for accelerated work and 8 GB+ for the default multi-stage workflow. Large optional engines can require 12 to 16 GB or more. Check the benchmarks and engine guide.

Why does a longer reference clip not always improve the clone?

Cloning is zero-shot: the clip is a prompt, not training data. Use 5 to 15 seconds of one speaker, close to the microphone, without music, noise, or reverb. Match the tone and pace you want in the output. For training, see data preparation and training.

Can I use generated audio commercially?

VoiceStudio's application license does not restrict generated audio, but it does not grant rights under a model's separate terms. The default OmniVoice repository labels its pretrained weights CC-BY-NC and includes a tokenizer under separate community terms. Review the selected model terms before commercial use.

Does VoiceStudio collect data?

Not unless you opt in. Analytics is off by default and skipping consent keeps it off. When enabled, the app sends allowlisted, content-free usage metadata. Text, audio, file names, voices, and projects are excluded. Change this at Settings → Privacy.

How do I remove VoiceStudio and its data?

Use scripts/uninstall.sh on macOS/Linux or scripts\uninstall.ps1 on Windows. Both show a dry run before deletion. See the uninstall guide for every path.

Community and contributing

Support development

VoiceStudio is free and has no paid tier. Donations fund development and infrastructure.

Ko-fi · PayPal · Sponsorship details

License

VoiceStudio is licensed under AGPL-3.0. You may run it, modify it, and use it internally. The application license itself does not restrict selling generated audio, but downloaded model and tokenizer terms may. If you modify VoiceStudio and provide that modified version as a network service, AGPL requires you to offer the corresponding source under the same license. A commercial license for VoiceStudio-owned code is available for proprietary embedding; it does not relicense third-party models. Contact VoiceStudio@palash.dev. See LICENSE-NOTICE.md for the plain-language scope.

Optional engines and downloaded models retain their own licenses. The bundled omnivoice/ Python code is Apache-2.0 upstream; the default downloaded weights and audio tokenizer use separate terms.

Acknowledgments

VoiceStudio builds on OmniVoice, WhisperX, Demucs, Pyannote, CTranslate2, AudioSeal, Tauri, Supertonic, Sherpa-ONNX, GPT-SoVITS, and PocketTTS.

About

VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

14.0k stars

Watchers

67 watching

Forks

Releases

Sponsor this project

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

1,893 Commits

Folders and files

NameName
Last commit message
Last commit date
VoiceStudio ranking on TrendshiftVoiceStudio logo

VoiceStudio

Previously OmniVoice-Studio

Clone voices, dub video, dictate, and produce long-form audio on your own hardware.

16 TTS engines · 11 ASR engines · 646-language catalogue · macOS, Windows, Linux, and Docker

No account, API key, subscription, or usage meter for the local workflow.

Install · Features · Compare · Requirements · Engines · Architecture · API · Docs · 简体中文

GitHub starsTotal downloadsLatest releaseAGPL-3.0 licenseDiscord community

Download VoiceStudio

Switching TTS engines from the VoiceStudio status bar

Warning

Active beta. Use the latest release for stable work. main contains the newest fixes and may change between releases. Report problems through GitHub Issues.

At a glance

VoiceStudio
WorkflowsVoice cloning and design, video dubbing, dictation, stories, audiobooks, batch generation
Language catalogue646 TTS languages; actual coverage and quality depend on the selected engine
Engines16 TTS · 11 ASR · switch in Model Catalogue or with Ctrl/Cmd+E
PlatformsmacOS 13.3+ on Apple Silicon · Windows 10/11 x64 · Linux x86_64 with glibc 2.39+
ComputeCUDA · Apple Silicon MPS/MLX · ROCm on Linux · CPU · optional remote workers
InterfacesDesktop app · local REST/SSE/WebSocket API · OpenAI-compatible audio API · MCP Server
StorageVoices, projects, settings, and outputs stay on the machine by default
LicenseAGPL-3.0 application; downloaded models keep their upstream terms

Install

Download a package from the latest release, then follow the platform guide.

PlatformPackageGuide
macOS 13.3+Apple Silicon DMGInstall on macOS
Windows 10/11x64 MSI; choose the current-user build when listed to install without admin accessInstall on Windows
LinuxAppImage, x86_64 with glibc 2.39+Install on Linux
DockerCUDA, ROCm, CPU, and worker-only GPU profilesRun with Docker

First launch creates a managed Python environment and downloads the default model. Later launches reuse both.

Note

On macOS, first launch needs a one-time right-click, then Open approval. Intel Macs cannot run the local Python backend; use a remote backend instead.

First voice

  1. Launch VoiceStudio and open Voice Cloning.
  2. Add a clean voice sample. Three seconds works; 5 to 15 seconds usually gives a better prompt.
  3. Enter text, choose a language, then select Generate.

Run from source

Install the development prerequisites, then:

git clone https://github.com/debpalash/VoiceStudio.git
cd VoiceStudio
bun install
bun run desktop

Use bun run dev for the browser UI. See Contributing for services, tests, and platform packages.

If setup fails

Features

AreaIncluded
Voice CloningZero-shot synthesis from a short reference clip
Voice DesignCreate a voice from age, accent, pitch, style, and delivery instructions
Video DubbingTranscribe, translate, preserve speakers, synthesize, and export video
Stories and audiobooksMulti-voice scripts · EPUB/PDF import · chapter rendering · .m4b export
Dictation WidgetSystem-wide shortcut, live transcription, optional local-LLM cleanup
Vocal IsolationDemucs speech/background separation
Speaker DiarizationPyannote and WhisperX speaker assignment
Batch QueueQueue large sets of audio and video jobs with per-job progress
Model CatalogueInstall, remove, select, and route TTS, ASR, and LLM models
Remote Model DownloadsInstall models on enrolled remote workers with live progress
GPU Auto-DetectCUDA, MPS, ROCm, and CPU routing with per-engine checks
AI WatermarkAudioSeal embedding and detection
MCP ServerSynthesis and transcription tools for MCP clients
DiagnosticsSelf-checks, error journal, logs, and scrubbed support bundles
Local-firstCore creation stays local; network-backed features are explicit opt-ins
ExtensibleRegistry-based TTS, ASR, and plugin interfaces
VoiceStudio Model CatalogueSaving a gallery voice as a local profile
Model Catalogue: engine, device, and install stateGallery: save a shared voice as a local profile

Comparison

VoiceStudio trades managed cloud compute for local control. This is the practical difference:

VoiceStudioTypical hosted voice service
Best fitPrivate, offline, self-hosted, or high-volume workFast setup without local model management
Data pathLocal by default; remote features are opt-inAudio and text are processed by the provider
Cost modelFree software; you supply the hardwareSubscription, credits, or metered API use
SetupInstall the app and model weightsCreate an account and use the web app or API
PerformanceDepends on your engine and hardwareProvider manages compute and scaling
Offline useYes, after required models are installedUsually requires a network connection
CustomizationSource, engines, models, API, and routing are openLimited to provider options
MaintenanceYou manage updates, disk, and computeProvider manages infrastructure

Requirements

Requirements vary by engine. These values cover the default local workflow.

MinimumRecommended
OSWindows 10 x64 · macOS 13.3 Apple Silicon · Linux x86_64 with glibc 2.39+Current supported OS release
RAM8 GB16 GB+
Disk10 GB free20 GB+ SSD
GPUOptional; CPU mode is supportedNVIDIA CUDA or Apple Silicon
VRAM4 GB when using a GPU8 GB+; large optional engines need more
Python from source3.11+3.11 or 3.12

ROCm is Linux-only and opt-in. Windows AMD/Ryzen AI uses CPU. Systems with limited VRAM offload work to CPU when required. See performance, benchmarks, and engine disk usage.

Engines

Engine support is capability-specific. Check cloning, language, platform, memory, and license before choosing one. Full setup guides: docs/engines.

Text to speech

EngineLanguagesCloneInstructLinuxmacOS ARMWindowsLicense
VoiceStudio (default, powered by k2-fsa/OmniVoice)600+YesYesCUDA/CPUMPSCUDA/CPUAGPL-3.0 app · Apache-2.0 code, CC-BY-NC weights³
CosyVoice 39 + 18 dialectsYesYesCUDA/CPUCPUCUDA/CPUApache-2.0
GPT-SoVITS5YesNoCUDA/CPUNoCUDA/CPUMIT
VoxCPM230YesYesCUDA/CPUMPSCUDA/CPUApache-2.0
MOSS-TTS-Nano20YesNoCUDA/CPUCPUCUDA/CPUApache-2.0
KittenTTSEnglishNoNoCPUCPUCPUMIT
MLX-AudioModel-dependentVariesVariesNoMLXNoVaries
Sherpa-ONNX20+NoNoCUDA/CPUCPUCUDA/CPUApache-2.0
IndexTTS 2.5ZH · EN · JA · ES · ARYesNoCUDA/CPUCPUCUDA/CPUBilibili model license¹
OmniVoice GGUF600+YesYesCUDA/CPUMPS/CPUCUDA/CPUAGPL-3.0 app · review the derivative model terms³
OmniVoice (subprocess)600+YesYesCUDA/CPUMPSCUDA/CPUAGPL-3.0 app · Apache-2.0 code, CC-BY-NC weights³
PocketTTSEN · FR · DE · PT · IT · ESYesNoCPUCPUCPUCC-BY-4.0, gated²
Supertonic 331NoNoCPUCPUCPUOpenRAIL-M
MOSS-TTS-v1.531YesNoCUDA/CPUCPUCUDA/CPUApache-2.0
dots.tts24YesNoCUDA/CPUCPUNoApache-2.0
Confucius4-TTS14YesNoCUDA/CPUCPUCUDA/CPUApache-2.0

⚡ Installed or registered on demand.

¹ IndexTTS 2.5 requires a separate written Bilibili license above 100 million monthly active users or RMB 1 billion annual revenue. Review the model license.

² PocketTTS shows its gated-access and CC-BY-4.0 terms before first use.

³ The OmniVoice snapshot also includes an audio tokenizer under separate Boson Higgs Audio 2 and Meta Llama community terms. VoiceStudio's application license does not replace model or tokenizer terms.

Clone-less engines cannot preserve a reference speaker in dubbing or pinned-voice batch jobs. VoiceStudio rejects those jobs instead of silently changing engines. Heavy engines have separate memory and platform limits; check their engine guide first.

Speech to text

EngineIDLanguagesBest fit
WhisperX (default)whisperx~100Dubbing, subtitles, word-level timing
Faster-Whisperfaster-whisper~100General cross-platform transcription
Faster-Whisper (isolated)faster-whisper-isolated~100Crash-isolated batch transcription
MLX Whispermlx-whisper~100Apple Silicon
PyTorch Whisperpytorch-whisper~100CUDA, MPS, and CPU fallback
Parakeet TDTnemo-parakeetEnglish + 25 EUFast CPU/CUDA transcription
Parakeet TDT v3 (MLX)parakeet-mlx25 EUApple Silicon dictation and word timestamps
MoonshinemoonshineEnglishLow-power, low-latency ONNX
FunASRfunasr50+VAD and inline diarization
sherpa-onnx (live dictation)sherpa-onnx-asrModel-dependentStreaming CPU dictation
OpenAI-compatible⚠️ configured serveropenai-compat-asrServer-dependentLocal gigastt/Qwen3-ASR or a remote endpoint; audio goes only to that server

WhisperX and Faster-Whisper retry with int8 when efficient float16 is unavailable. Pin ASR_COMPUTE_TYPE=int8 or float32 only if automatic selection still fails.

Architecture

Tauri v2 desktop shell (Rust)
│ IPC
React + Vite UI
│ HTTP · SSE · WebSocket on localhost:3900
FastAPI backend
├── TTS / ASR engine registries
├── dubbing / audio / long-form pipelines
├── OpenAI-compatible API and MCP server
└── SQLite + Alembic → omnivoice_data/
LayerPathResponsibility
Desktop shellfrontend/src-tauri/Window lifecycle, tray, shortcuts, updater, sidecar bootstrap
Frontendfrontend/src/React UI, Zustand state, API and event clients, i18n
APIbackend/api/REST routes, schemas, auth boundaries, streaming
Core servicesbackend/services/Generation, dubbing, audio processing, persistence
Enginesbackend/engines/Isolated and optional engine adapters
Worker systembackend/worker/Authenticated remote compute and job transport
Dataomnivoice_data/Projects, voices, settings, logs, and SQLite state
Deliveryscripts/, deploy/, .github/workflows/Development, packaging, containers, releases, CI

Network boundary

  • The desktop talks to a loopback-only backend on localhost:3900.
  • Loopback API calls need no server key. Remote access requires a share PIN or API key.
  • Remote workers and OpenAI-compatible ASR are opt-in. Loopback ASR may use HTTP and keeps audio on the machine; non-loopback endpoints require HTTPS, and redirects are not followed.
  • Analytics is off until consent. If enabled, it sends allowlisted, content-free usage metadata. It never sends text, audio, file names, or projects.

Local speech platform and OpenAI-compatible API

Point an OpenAI-compatible audio client at the local backend:

- base_url="https://api.openai.com/v1"+ base_url="http://localhost:3900/v1"
EndpointPurpose
POST /v1/audio/speechTTS to mp3, opus, aac, flac, wav, or pcm; select a profile with voice and an engine with model
POST /v1/audio/transcriptionsSTT to json, text, verbose_json, srt, or vtt
WS /v1/audio/transcriptions/streamLive PCM/WebM transcription with partial, utterance, and session-final events
GET /.well-known/voicestudio-speechDiscover HTTP, WebSocket, MCP, and native dictation-control transports
GET /v1/audio/voicesList local voice profiles and engines
fromopenaiimportOpenAIclient=OpenAI(base_url="http://localhost:3900/v1", api_key="local")
withclient.audio.speech.with_streaming_response.create(
model="tts-1",
voice="<profile-id>",
input="Made on my own hardware.",
response_format="wav",
) asresponse:
response.stream_to_file("speech.wav")

The bundled Rust control sidecar lets Herdr, coding agents, VS Code, desktop apps, and TUIs trigger the system-wide dictation flow or reuse its native text insertion. See the speech platform guide. The full API reference is in Settings → OpenAPI Reference. For LAN, Tailscale, or proxy access, read API authentication before exposing the backend.

Agent skills

Install the VoiceStudio skills for Claude Code, Codex, Cursor, and other skills.sh-compatible agents:

npx skills add debpalash/VoiceStudio
  • omnivoice: synthesize speech and transcribe audio through local VoiceStudio.
  • oss-maintainer: the repository's open-source maintenance workflow.

Google Colab

Open in Colab

The notebook runs the app and web UI on a Colab GPU. Colab is remote compute, so uploaded audio and project data do not remain local to your machine.

Documentation

NeedRead
InstallmacOS · Windows · Linux · Docker
Fix setupTroubleshooting · model downloads · Hugging Face token
Choose an engineEngine guides · benchmarks · expressive speech
Tune hardwarePerformance · remote workers
Build integrationsSpeech platform · Private production API · API auth · MCP · examples
Build VoiceStudioContributing · engine acceptance
Track changesChangelog · roadmap · latest release
Remove everythingUninstall guide

FAQ

Does it work on Apple Silicon and Intel Macs?

Apple Silicon is supported with MPS and MLX options. Intel Macs cannot run the local backend because current PyTorch wheels are unavailable; they can connect to a remote backend. See macOS installation.

How much VRAM do I need?

A GPU is optional. Use 4 GB VRAM as the minimum for accelerated work and 8 GB+ for the default multi-stage workflow. Large optional engines can require 12 to 16 GB or more. Check the benchmarks and engine guide.

Why does a longer reference clip not always improve the clone?

Cloning is zero-shot: the clip is a prompt, not training data. Use 5 to 15 seconds of one speaker, close to the microphone, without music, noise, or reverb. Match the tone and pace you want in the output. For training, see data preparation and training.

Can I use generated audio commercially?

VoiceStudio's application license does not restrict generated audio, but it does not grant rights under a model's separate terms. The default OmniVoice repository labels its pretrained weights CC-BY-NC and includes a tokenizer under separate community terms. Review the selected model terms before commercial use.

Does VoiceStudio collect data?

Not unless you opt in. Analytics is off by default and skipping consent keeps it off. When enabled, the app sends allowlisted, content-free usage metadata. Text, audio, file names, voices, and projects are excluded. Change this at Settings → Privacy.

How do I remove VoiceStudio and its data?

Use scripts/uninstall.sh on macOS/Linux or scripts\uninstall.ps1 on Windows. Both show a dry run before deletion. See the uninstall guide for every path.

Community and contributing

Support development

VoiceStudio is free and has no paid tier. Donations fund development and infrastructure.

Ko-fi · PayPal · Sponsorship details

License

VoiceStudio is licensed under AGPL-3.0. You may run it, modify it, and use it internally. The application license itself does not restrict selling generated audio, but downloaded model and tokenizer terms may. If you modify VoiceStudio and provide that modified version as a network service, AGPL requires you to offer the corresponding source under the same license. A commercial license for VoiceStudio-owned code is available for proprietary embedding; it does not relicense third-party models. Contact VoiceStudio@palash.dev. See LICENSE-NOTICE.md for the plain-language scope.

Optional engines and downloaded models retain their own licenses. The bundled omnivoice/ Python code is Apache-2.0 upstream; the default downloaded weights and audio tokenizer use separate terms.

Acknowledgments

VoiceStudio builds on OmniVoice, WhisperX, Demucs, Pyannote, CTranslate2, AudioSeal, Tauri, Supertonic, Sherpa-ONNX, GPT-SoVITS, and PocketTTS.

About

VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

14.0k stars

Watchers

67 watching

Forks

Releases

Sponsor this project

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

1,893 Commits

Folders and files

NameName
Last commit message
Last commit date
VoiceStudio ranking on TrendshiftVoiceStudio logo

VoiceStudio

Previously OmniVoice-Studio

Clone voices, dub video, dictate, and produce long-form audio on your own hardware.

16 TTS engines · 11 ASR engines · 646-language catalogue · macOS, Windows, Linux, and Docker

No account, API key, subscription, or usage meter for the local workflow.

Install · Features · Compare · Requirements · Engines · Architecture · API · Docs · 简体中文

GitHub starsTotal downloadsLatest releaseAGPL-3.0 licenseDiscord community

Download VoiceStudio

Switching TTS engines from the VoiceStudio status bar

Warning

Active beta. Use the latest release for stable work. main contains the newest fixes and may change between releases. Report problems through GitHub Issues.

At a glance

VoiceStudio
WorkflowsVoice cloning and design, video dubbing, dictation, stories, audiobooks, batch generation
Language catalogue646 TTS languages; actual coverage and quality depend on the selected engine
Engines16 TTS · 11 ASR · switch in Model Catalogue or with Ctrl/Cmd+E
PlatformsmacOS 13.3+ on Apple Silicon · Windows 10/11 x64 · Linux x86_64 with glibc 2.39+
ComputeCUDA · Apple Silicon MPS/MLX · ROCm on Linux · CPU · optional remote workers
InterfacesDesktop app · local REST/SSE/WebSocket API · OpenAI-compatible audio API · MCP Server
StorageVoices, projects, settings, and outputs stay on the machine by default
LicenseAGPL-3.0 application; downloaded models keep their upstream terms

Install

Download a package from the latest release, then follow the platform guide.

PlatformPackageGuide
macOS 13.3+Apple Silicon DMGInstall on macOS
Windows 10/11x64 MSI; choose the current-user build when listed to install without admin accessInstall on Windows
LinuxAppImage, x86_64 with glibc 2.39+Install on Linux
DockerCUDA, ROCm, CPU, and worker-only GPU profilesRun with Docker

First launch creates a managed Python environment and downloads the default model. Later launches reuse both.

Note

On macOS, first launch needs a one-time right-click, then Open approval. Intel Macs cannot run the local Python backend; use a remote backend instead.

First voice

  1. Launch VoiceStudio and open Voice Cloning.
  2. Add a clean voice sample. Three seconds works; 5 to 15 seconds usually gives a better prompt.
  3. Enter text, choose a language, then select Generate.

Run from source

Install the development prerequisites, then:

git clone https://github.com/debpalash/VoiceStudio.git
cd VoiceStudio
bun install
bun run desktop

Use bun run dev for the browser UI. See Contributing for services, tests, and platform packages.

If setup fails

Features

AreaIncluded
Voice CloningZero-shot synthesis from a short reference clip
Voice DesignCreate a voice from age, accent, pitch, style, and delivery instructions
Video DubbingTranscribe, translate, preserve speakers, synthesize, and export video
Stories and audiobooksMulti-voice scripts · EPUB/PDF import · chapter rendering · .m4b export
Dictation WidgetSystem-wide shortcut, live transcription, optional local-LLM cleanup
Vocal IsolationDemucs speech/background separation
Speaker DiarizationPyannote and WhisperX speaker assignment
Batch QueueQueue large sets of audio and video jobs with per-job progress
Model CatalogueInstall, remove, select, and route TTS, ASR, and LLM models
Remote Model DownloadsInstall models on enrolled remote workers with live progress
GPU Auto-DetectCUDA, MPS, ROCm, and CPU routing with per-engine checks
AI WatermarkAudioSeal embedding and detection
MCP ServerSynthesis and transcription tools for MCP clients
DiagnosticsSelf-checks, error journal, logs, and scrubbed support bundles
Local-firstCore creation stays local; network-backed features are explicit opt-ins
ExtensibleRegistry-based TTS, ASR, and plugin interfaces
VoiceStudio Model CatalogueSaving a gallery voice as a local profile
Model Catalogue: engine, device, and install stateGallery: save a shared voice as a local profile

Comparison

VoiceStudio trades managed cloud compute for local control. This is the practical difference:

VoiceStudioTypical hosted voice service
Best fitPrivate, offline, self-hosted, or high-volume workFast setup without local model management
Data pathLocal by default; remote features are opt-inAudio and text are processed by the provider
Cost modelFree software; you supply the hardwareSubscription, credits, or metered API use
SetupInstall the app and model weightsCreate an account and use the web app or API
PerformanceDepends on your engine and hardwareProvider manages compute and scaling
Offline useYes, after required models are installedUsually requires a network connection
CustomizationSource, engines, models, API, and routing are openLimited to provider options
MaintenanceYou manage updates, disk, and computeProvider manages infrastructure

Requirements

Requirements vary by engine. These values cover the default local workflow.

MinimumRecommended
OSWindows 10 x64 · macOS 13.3 Apple Silicon · Linux x86_64 with glibc 2.39+Current supported OS release
RAM8 GB16 GB+
Disk10 GB free20 GB+ SSD
GPUOptional; CPU mode is supportedNVIDIA CUDA or Apple Silicon
VRAM4 GB when using a GPU8 GB+; large optional engines need more
Python from source3.11+3.11 or 3.12

ROCm is Linux-only and opt-in. Windows AMD/Ryzen AI uses CPU. Systems with limited VRAM offload work to CPU when required. See performance, benchmarks, and engine disk usage.

Engines

Engine support is capability-specific. Check cloning, language, platform, memory, and license before choosing one. Full setup guides: docs/engines.

Text to speech

EngineLanguagesCloneInstructLinuxmacOS ARMWindowsLicense
VoiceStudio (default, powered by k2-fsa/OmniVoice)600+YesYesCUDA/CPUMPSCUDA/CPUAGPL-3.0 app · Apache-2.0 code, CC-BY-NC weights³
CosyVoice 39 + 18 dialectsYesYesCUDA/CPUCPUCUDA/CPUApache-2.0
GPT-SoVITS5YesNoCUDA/CPUNoCUDA/CPUMIT
VoxCPM230YesYesCUDA/CPUMPSCUDA/CPUApache-2.0
MOSS-TTS-Nano20YesNoCUDA/CPUCPUCUDA/CPUApache-2.0
KittenTTSEnglishNoNoCPUCPUCPUMIT
MLX-AudioModel-dependentVariesVariesNoMLXNoVaries
Sherpa-ONNX20+NoNoCUDA/CPUCPUCUDA/CPUApache-2.0
IndexTTS 2.5ZH · EN · JA · ES · ARYesNoCUDA/CPUCPUCUDA/CPUBilibili model license¹
OmniVoice GGUF600+YesYesCUDA/CPUMPS/CPUCUDA/CPUAGPL-3.0 app · review the derivative model terms³
OmniVoice (subprocess)600+YesYesCUDA/CPUMPSCUDA/CPUAGPL-3.0 app · Apache-2.0 code, CC-BY-NC weights³
PocketTTSEN · FR · DE · PT · IT · ESYesNoCPUCPUCPUCC-BY-4.0, gated²
Supertonic 331NoNoCPUCPUCPUOpenRAIL-M
MOSS-TTS-v1.531YesNoCUDA/CPUCPUCUDA/CPUApache-2.0
dots.tts24YesNoCUDA/CPUCPUNoApache-2.0
Confucius4-TTS14YesNoCUDA/CPUCPUCUDA/CPUApache-2.0

⚡ Installed or registered on demand.

¹ IndexTTS 2.5 requires a separate written Bilibili license above 100 million monthly active users or RMB 1 billion annual revenue. Review the model license.

² PocketTTS shows its gated-access and CC-BY-4.0 terms before first use.

³ The OmniVoice snapshot also includes an audio tokenizer under separate Boson Higgs Audio 2 and Meta Llama community terms. VoiceStudio's application license does not replace model or tokenizer terms.

Clone-less engines cannot preserve a reference speaker in dubbing or pinned-voice batch jobs. VoiceStudio rejects those jobs instead of silently changing engines. Heavy engines have separate memory and platform limits; check their engine guide first.

Speech to text

EngineIDLanguagesBest fit
WhisperX (default)whisperx~100Dubbing, subtitles, word-level timing
Faster-Whisperfaster-whisper~100General cross-platform transcription
Faster-Whisper (isolated)faster-whisper-isolated~100Crash-isolated batch transcription
MLX Whispermlx-whisper~100Apple Silicon
PyTorch Whisperpytorch-whisper~100CUDA, MPS, and CPU fallback
Parakeet TDTnemo-parakeetEnglish + 25 EUFast CPU/CUDA transcription
Parakeet TDT v3 (MLX)parakeet-mlx25 EUApple Silicon dictation and word timestamps
MoonshinemoonshineEnglishLow-power, low-latency ONNX
FunASRfunasr50+VAD and inline diarization
sherpa-onnx (live dictation)sherpa-onnx-asrModel-dependentStreaming CPU dictation
OpenAI-compatible⚠️ configured serveropenai-compat-asrServer-dependentLocal gigastt/Qwen3-ASR or a remote endpoint; audio goes only to that server

WhisperX and Faster-Whisper retry with int8 when efficient float16 is unavailable. Pin ASR_COMPUTE_TYPE=int8 or float32 only if automatic selection still fails.

Architecture

Tauri v2 desktop shell (Rust)
│ IPC
React + Vite UI
│ HTTP · SSE · WebSocket on localhost:3900
FastAPI backend
├── TTS / ASR engine registries
├── dubbing / audio / long-form pipelines
├── OpenAI-compatible API and MCP server
└── SQLite + Alembic → omnivoice_data/
LayerPathResponsibility
Desktop shellfrontend/src-tauri/Window lifecycle, tray, shortcuts, updater, sidecar bootstrap
Frontendfrontend/src/React UI, Zustand state, API and event clients, i18n
APIbackend/api/REST routes, schemas, auth boundaries, streaming
Core servicesbackend/services/Generation, dubbing, audio processing, persistence
Enginesbackend/engines/Isolated and optional engine adapters
Worker systembackend/worker/Authenticated remote compute and job transport
Dataomnivoice_data/Projects, voices, settings, logs, and SQLite state
Deliveryscripts/, deploy/, .github/workflows/Development, packaging, containers, releases, CI

Network boundary

  • The desktop talks to a loopback-only backend on localhost:3900.
  • Loopback API calls need no server key. Remote access requires a share PIN or API key.
  • Remote workers and OpenAI-compatible ASR are opt-in. Loopback ASR may use HTTP and keeps audio on the machine; non-loopback endpoints require HTTPS, and redirects are not followed.
  • Analytics is off until consent. If enabled, it sends allowlisted, content-free usage metadata. It never sends text, audio, file names, or projects.

Local speech platform and OpenAI-compatible API

Point an OpenAI-compatible audio client at the local backend:

- base_url="https://api.openai.com/v1"+ base_url="http://localhost:3900/v1"
EndpointPurpose
POST /v1/audio/speechTTS to mp3, opus, aac, flac, wav, or pcm; select a profile with voice and an engine with model
POST /v1/audio/transcriptionsSTT to json, text, verbose_json, srt, or vtt
WS /v1/audio/transcriptions/streamLive PCM/WebM transcription with partial, utterance, and session-final events
GET /.well-known/voicestudio-speechDiscover HTTP, WebSocket, MCP, and native dictation-control transports
GET /v1/audio/voicesList local voice profiles and engines
fromopenaiimportOpenAIclient=OpenAI(base_url="http://localhost:3900/v1", api_key="local")
withclient.audio.speech.with_streaming_response.create(
model="tts-1",
voice="<profile-id>",
input="Made on my own hardware.",
response_format="wav",
) asresponse:
response.stream_to_file("speech.wav")

The bundled Rust control sidecar lets Herdr, coding agents, VS Code, desktop apps, and TUIs trigger the system-wide dictation flow or reuse its native text insertion. See the speech platform guide. The full API reference is in Settings → OpenAPI Reference. For LAN, Tailscale, or proxy access, read API authentication before exposing the backend.

Agent skills

Install the VoiceStudio skills for Claude Code, Codex, Cursor, and other skills.sh-compatible agents:

npx skills add debpalash/VoiceStudio
  • omnivoice: synthesize speech and transcribe audio through local VoiceStudio.
  • oss-maintainer: the repository's open-source maintenance workflow.

Google Colab

Open in Colab

The notebook runs the app and web UI on a Colab GPU. Colab is remote compute, so uploaded audio and project data do not remain local to your machine.

Documentation

NeedRead
InstallmacOS · Windows · Linux · Docker
Fix setupTroubleshooting · model downloads · Hugging Face token
Choose an engineEngine guides · benchmarks · expressive speech
Tune hardwarePerformance · remote workers
Build integrationsSpeech platform · Private production API · API auth · MCP · examples
Build VoiceStudioContributing · engine acceptance
Track changesChangelog · roadmap · latest release
Remove everythingUninstall guide

FAQ

Does it work on Apple Silicon and Intel Macs?

Apple Silicon is supported with MPS and MLX options. Intel Macs cannot run the local backend because current PyTorch wheels are unavailable; they can connect to a remote backend. See macOS installation.

How much VRAM do I need?

A GPU is optional. Use 4 GB VRAM as the minimum for accelerated work and 8 GB+ for the default multi-stage workflow. Large optional engines can require 12 to 16 GB or more. Check the benchmarks and engine guide.

Why does a longer reference clip not always improve the clone?

Cloning is zero-shot: the clip is a prompt, not training data. Use 5 to 15 seconds of one speaker, close to the microphone, without music, noise, or reverb. Match the tone and pace you want in the output. For training, see data preparation and training.

Can I use generated audio commercially?

VoiceStudio's application license does not restrict generated audio, but it does not grant rights under a model's separate terms. The default OmniVoice repository labels its pretrained weights CC-BY-NC and includes a tokenizer under separate community terms. Review the selected model terms before commercial use.

Does VoiceStudio collect data?

Not unless you opt in. Analytics is off by default and skipping consent keeps it off. When enabled, the app sends allowlisted, content-free usage metadata. Text, audio, file names, voices, and projects are excluded. Change this at Settings → Privacy.

How do I remove VoiceStudio and its data?

Use scripts/uninstall.sh on macOS/Linux or scripts\uninstall.ps1 on Windows. Both show a dry run before deletion. See the uninstall guide for every path.

Community and contributing

Support development

VoiceStudio is free and has no paid tier. Donations fund development and infrastructure.

Ko-fi · PayPal · Sponsorship details

License

VoiceStudio is licensed under AGPL-3.0. You may run it, modify it, and use it internally. The application license itself does not restrict selling generated audio, but downloaded model and tokenizer terms may. If you modify VoiceStudio and provide that modified version as a network service, AGPL requires you to offer the corresponding source under the same license. A commercial license for VoiceStudio-owned code is available for proprietary embedding; it does not relicense third-party models. Contact VoiceStudio@palash.dev. See LICENSE-NOTICE.md for the plain-language scope.

Optional engines and downloaded models retain their own licenses. The bundled omnivoice/ Python code is Apache-2.0 upstream; the default downloaded weights and audio tokenizer use separate terms.

Acknowledgments

VoiceStudio builds on OmniVoice, WhisperX, Demucs, Pyannote, CTranslate2, AudioSeal, Tauri, Supertonic, Sherpa-ONNX, GPT-SoVITS, and PocketTTS.

About

VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

14.0k stars

Watchers

67 watching

Forks

Releases

Sponsor this project

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

1,893 Commits

Folders and files

NameName
Last commit message
Last commit date
VoiceStudio ranking on TrendshiftVoiceStudio logo

VoiceStudio

Previously OmniVoice-Studio

Clone voices, dub video, dictate, and produce long-form audio on your own hardware.

16 TTS engines · 11 ASR engines · 646-language catalogue · macOS, Windows, Linux, and Docker

No account, API key, subscription, or usage meter for the local workflow.

Install · Features · Compare · Requirements · Engines · Architecture · API · Docs · 简体中文

GitHub starsTotal downloadsLatest releaseAGPL-3.0 licenseDiscord community

Download VoiceStudio

Switching TTS engines from the VoiceStudio status bar

Warning

Active beta. Use the latest release for stable work. main contains the newest fixes and may change between releases. Report problems through GitHub Issues.

At a glance

VoiceStudio
WorkflowsVoice cloning and design, video dubbing, dictation, stories, audiobooks, batch generation
Language catalogue646 TTS languages; actual coverage and quality depend on the selected engine
Engines16 TTS · 11 ASR · switch in Model Catalogue or with Ctrl/Cmd+E
PlatformsmacOS 13.3+ on Apple Silicon · Windows 10/11 x64 · Linux x86_64 with glibc 2.39+
ComputeCUDA · Apple Silicon MPS/MLX · ROCm on Linux · CPU · optional remote workers
InterfacesDesktop app · local REST/SSE/WebSocket API · OpenAI-compatible audio API · MCP Server
StorageVoices, projects, settings, and outputs stay on the machine by default
LicenseAGPL-3.0 application; downloaded models keep their upstream terms

Install

Download a package from the latest release, then follow the platform guide.

PlatformPackageGuide
macOS 13.3+Apple Silicon DMGInstall on macOS
Windows 10/11x64 MSI; choose the current-user build when listed to install without admin accessInstall on Windows
LinuxAppImage, x86_64 with glibc 2.39+Install on Linux
DockerCUDA, ROCm, CPU, and worker-only GPU profilesRun with Docker

First launch creates a managed Python environment and downloads the default model. Later launches reuse both.

Note

On macOS, first launch needs a one-time right-click, then Open approval. Intel Macs cannot run the local Python backend; use a remote backend instead.

First voice

  1. Launch VoiceStudio and open Voice Cloning.
  2. Add a clean voice sample. Three seconds works; 5 to 15 seconds usually gives a better prompt.
  3. Enter text, choose a language, then select Generate.

Run from source

Install the development prerequisites, then:

git clone https://github.com/debpalash/VoiceStudio.git
cd VoiceStudio
bun install
bun run desktop

Use bun run dev for the browser UI. See Contributing for services, tests, and platform packages.

If setup fails

Features

AreaIncluded
Voice CloningZero-shot synthesis from a short reference clip
Voice DesignCreate a voice from age, accent, pitch, style, and delivery instructions
Video DubbingTranscribe, translate, preserve speakers, synthesize, and export video
Stories and audiobooksMulti-voice scripts · EPUB/PDF import · chapter rendering · .m4b export
Dictation WidgetSystem-wide shortcut, live transcription, optional local-LLM cleanup
Vocal IsolationDemucs speech/background separation
Speaker DiarizationPyannote and WhisperX speaker assignment
Batch QueueQueue large sets of audio and video jobs with per-job progress
Model CatalogueInstall, remove, select, and route TTS, ASR, and LLM models
Remote Model DownloadsInstall models on enrolled remote workers with live progress
GPU Auto-DetectCUDA, MPS, ROCm, and CPU routing with per-engine checks
AI WatermarkAudioSeal embedding and detection
MCP ServerSynthesis and transcription tools for MCP clients
DiagnosticsSelf-checks, error journal, logs, and scrubbed support bundles
Local-firstCore creation stays local; network-backed features are explicit opt-ins
ExtensibleRegistry-based TTS, ASR, and plugin interfaces
VoiceStudio Model CatalogueSaving a gallery voice as a local profile
Model Catalogue: engine, device, and install stateGallery: save a shared voice as a local profile

Comparison

VoiceStudio trades managed cloud compute for local control. This is the practical difference:

VoiceStudioTypical hosted voice service
Best fitPrivate, offline, self-hosted, or high-volume workFast setup without local model management
Data pathLocal by default; remote features are opt-inAudio and text are processed by the provider
Cost modelFree software; you supply the hardwareSubscription, credits, or metered API use
SetupInstall the app and model weightsCreate an account and use the web app or API
PerformanceDepends on your engine and hardwareProvider manages compute and scaling
Offline useYes, after required models are installedUsually requires a network connection
CustomizationSource, engines, models, API, and routing are openLimited to provider options
MaintenanceYou manage updates, disk, and computeProvider manages infrastructure

Requirements

Requirements vary by engine. These values cover the default local workflow.

MinimumRecommended
OSWindows 10 x64 · macOS 13.3 Apple Silicon · Linux x86_64 with glibc 2.39+Current supported OS release
RAM8 GB16 GB+
Disk10 GB free20 GB+ SSD
GPUOptional; CPU mode is supportedNVIDIA CUDA or Apple Silicon
VRAM4 GB when using a GPU8 GB+; large optional engines need more
Python from source3.11+3.11 or 3.12

ROCm is Linux-only and opt-in. Windows AMD/Ryzen AI uses CPU. Systems with limited VRAM offload work to CPU when required. See performance, benchmarks, and engine disk usage.

Engines

Engine support is capability-specific. Check cloning, language, platform, memory, and license before choosing one. Full setup guides: docs/engines.

Text to speech

EngineLanguagesCloneInstructLinuxmacOS ARMWindowsLicense
VoiceStudio (default, powered by k2-fsa/OmniVoice)600+YesYesCUDA/CPUMPSCUDA/CPUAGPL-3.0 app · Apache-2.0 code, CC-BY-NC weights³
CosyVoice 39 + 18 dialectsYesYesCUDA/CPUCPUCUDA/CPUApache-2.0
GPT-SoVITS5YesNoCUDA/CPUNoCUDA/CPUMIT
VoxCPM230YesYesCUDA/CPUMPSCUDA/CPUApache-2.0
MOSS-TTS-Nano20YesNoCUDA/CPUCPUCUDA/CPUApache-2.0
KittenTTSEnglishNoNoCPUCPUCPUMIT
MLX-AudioModel-dependentVariesVariesNoMLXNoVaries
Sherpa-ONNX20+NoNoCUDA/CPUCPUCUDA/CPUApache-2.0
IndexTTS 2.5ZH · EN · JA · ES · ARYesNoCUDA/CPUCPUCUDA/CPUBilibili model license¹
OmniVoice GGUF600+YesYesCUDA/CPUMPS/CPUCUDA/CPUAGPL-3.0 app · review the derivative model terms³
OmniVoice (subprocess)600+YesYesCUDA/CPUMPSCUDA/CPUAGPL-3.0 app · Apache-2.0 code, CC-BY-NC weights³
PocketTTSEN · FR · DE · PT · IT · ESYesNoCPUCPUCPUCC-BY-4.0, gated²
Supertonic 331NoNoCPUCPUCPUOpenRAIL-M
MOSS-TTS-v1.531YesNoCUDA/CPUCPUCUDA/CPUApache-2.0
dots.tts24YesNoCUDA/CPUCPUNoApache-2.0
Confucius4-TTS14YesNoCUDA/CPUCPUCUDA/CPUApache-2.0

⚡ Installed or registered on demand.

¹ IndexTTS 2.5 requires a separate written Bilibili license above 100 million monthly active users or RMB 1 billion annual revenue. Review the model license.

² PocketTTS shows its gated-access and CC-BY-4.0 terms before first use.

³ The OmniVoice snapshot also includes an audio tokenizer under separate Boson Higgs Audio 2 and Meta Llama community terms. VoiceStudio's application license does not replace model or tokenizer terms.

Clone-less engines cannot preserve a reference speaker in dubbing or pinned-voice batch jobs. VoiceStudio rejects those jobs instead of silently changing engines. Heavy engines have separate memory and platform limits; check their engine guide first.

Speech to text

EngineIDLanguagesBest fit
WhisperX (default)whisperx~100Dubbing, subtitles, word-level timing
Faster-Whisperfaster-whisper~100General cross-platform transcription
Faster-Whisper (isolated)faster-whisper-isolated~100Crash-isolated batch transcription
MLX Whispermlx-whisper~100Apple Silicon
PyTorch Whisperpytorch-whisper~100CUDA, MPS, and CPU fallback
Parakeet TDTnemo-parakeetEnglish + 25 EUFast CPU/CUDA transcription
Parakeet TDT v3 (MLX)parakeet-mlx25 EUApple Silicon dictation and word timestamps
MoonshinemoonshineEnglishLow-power, low-latency ONNX
FunASRfunasr50+VAD and inline diarization
sherpa-onnx (live dictation)sherpa-onnx-asrModel-dependentStreaming CPU dictation
OpenAI-compatible⚠️ configured serveropenai-compat-asrServer-dependentLocal gigastt/Qwen3-ASR or a remote endpoint; audio goes only to that server

WhisperX and Faster-Whisper retry with int8 when efficient float16 is unavailable. Pin ASR_COMPUTE_TYPE=int8 or float32 only if automatic selection still fails.

Architecture

Tauri v2 desktop shell (Rust)
│ IPC
React + Vite UI
│ HTTP · SSE · WebSocket on localhost:3900
FastAPI backend
├── TTS / ASR engine registries
├── dubbing / audio / long-form pipelines
├── OpenAI-compatible API and MCP server
└── SQLite + Alembic → omnivoice_data/
LayerPathResponsibility
Desktop shellfrontend/src-tauri/Window lifecycle, tray, shortcuts, updater, sidecar bootstrap
Frontendfrontend/src/React UI, Zustand state, API and event clients, i18n
APIbackend/api/REST routes, schemas, auth boundaries, streaming
Core servicesbackend/services/Generation, dubbing, audio processing, persistence
Enginesbackend/engines/Isolated and optional engine adapters
Worker systembackend/worker/Authenticated remote compute and job transport
Dataomnivoice_data/Projects, voices, settings, logs, and SQLite state
Deliveryscripts/, deploy/, .github/workflows/Development, packaging, containers, releases, CI

Network boundary

  • The desktop talks to a loopback-only backend on localhost:3900.
  • Loopback API calls need no server key. Remote access requires a share PIN or API key.
  • Remote workers and OpenAI-compatible ASR are opt-in. Loopback ASR may use HTTP and keeps audio on the machine; non-loopback endpoints require HTTPS, and redirects are not followed.
  • Analytics is off until consent. If enabled, it sends allowlisted, content-free usage metadata. It never sends text, audio, file names, or projects.

Local speech platform and OpenAI-compatible API

Point an OpenAI-compatible audio client at the local backend:

- base_url="https://api.openai.com/v1"+ base_url="http://localhost:3900/v1"
EndpointPurpose
POST /v1/audio/speechTTS to mp3, opus, aac, flac, wav, or pcm; select a profile with voice and an engine with model
POST /v1/audio/transcriptionsSTT to json, text, verbose_json, srt, or vtt
WS /v1/audio/transcriptions/streamLive PCM/WebM transcription with partial, utterance, and session-final events
GET /.well-known/voicestudio-speechDiscover HTTP, WebSocket, MCP, and native dictation-control transports
GET /v1/audio/voicesList local voice profiles and engines
fromopenaiimportOpenAIclient=OpenAI(base_url="http://localhost:3900/v1", api_key="local")
withclient.audio.speech.with_streaming_response.create(
model="tts-1",
voice="<profile-id>",
input="Made on my own hardware.",
response_format="wav",
) asresponse:
response.stream_to_file("speech.wav")

The bundled Rust control sidecar lets Herdr, coding agents, VS Code, desktop apps, and TUIs trigger the system-wide dictation flow or reuse its native text insertion. See the speech platform guide. The full API reference is in Settings → OpenAPI Reference. For LAN, Tailscale, or proxy access, read API authentication before exposing the backend.

Agent skills

Install the VoiceStudio skills for Claude Code, Codex, Cursor, and other skills.sh-compatible agents:

npx skills add debpalash/VoiceStudio
  • omnivoice: synthesize speech and transcribe audio through local VoiceStudio.
  • oss-maintainer: the repository's open-source maintenance workflow.

Google Colab

Open in Colab

The notebook runs the app and web UI on a Colab GPU. Colab is remote compute, so uploaded audio and project data do not remain local to your machine.

Documentation

NeedRead
InstallmacOS · Windows · Linux · Docker
Fix setupTroubleshooting · model downloads · Hugging Face token
Choose an engineEngine guides · benchmarks · expressive speech
Tune hardwarePerformance · remote workers
Build integrationsSpeech platform · Private production API · API auth · MCP · examples
Build VoiceStudioContributing · engine acceptance
Track changesChangelog · roadmap · latest release
Remove everythingUninstall guide

FAQ

Does it work on Apple Silicon and Intel Macs?

Apple Silicon is supported with MPS and MLX options. Intel Macs cannot run the local backend because current PyTorch wheels are unavailable; they can connect to a remote backend. See macOS installation.

How much VRAM do I need?

A GPU is optional. Use 4 GB VRAM as the minimum for accelerated work and 8 GB+ for the default multi-stage workflow. Large optional engines can require 12 to 16 GB or more. Check the benchmarks and engine guide.

Why does a longer reference clip not always improve the clone?

Cloning is zero-shot: the clip is a prompt, not training data. Use 5 to 15 seconds of one speaker, close to the microphone, without music, noise, or reverb. Match the tone and pace you want in the output. For training, see data preparation and training.

Can I use generated audio commercially?

VoiceStudio's application license does not restrict generated audio, but it does not grant rights under a model's separate terms. The default OmniVoice repository labels its pretrained weights CC-BY-NC and includes a tokenizer under separate community terms. Review the selected model terms before commercial use.

Does VoiceStudio collect data?

Not unless you opt in. Analytics is off by default and skipping consent keeps it off. When enabled, the app sends allowlisted, content-free usage metadata. Text, audio, file names, voices, and projects are excluded. Change this at Settings → Privacy.

How do I remove VoiceStudio and its data?

Use scripts/uninstall.sh on macOS/Linux or scripts\uninstall.ps1 on Windows. Both show a dry run before deletion. See the uninstall guide for every path.

Community and contributing

Support development

VoiceStudio is free and has no paid tier. Donations fund development and infrastructure.

Ko-fi · PayPal · Sponsorship details

License

VoiceStudio is licensed under AGPL-3.0. You may run it, modify it, and use it internally. The application license itself does not restrict selling generated audio, but downloaded model and tokenizer terms may. If you modify VoiceStudio and provide that modified version as a network service, AGPL requires you to offer the corresponding source under the same license. A commercial license for VoiceStudio-owned code is available for proprietary embedding; it does not relicense third-party models. Contact VoiceStudio@palash.dev. See LICENSE-NOTICE.md for the plain-language scope.

Optional engines and downloaded models retain their own licenses. The bundled omnivoice/ Python code is Apache-2.0 upstream; the default downloaded weights and audio tokenizer use separate terms.

Acknowledgments

VoiceStudio builds on OmniVoice, WhisperX, Demucs, Pyannote, CTranslate2, AudioSeal, Tauri, Supertonic, Sherpa-ONNX, GPT-SoVITS, and PocketTTS.

About

VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

14.0k stars

Watchers

67 watching

Forks

Releases

Sponsor this project

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

1,893 Commits

Folders and files

NameName
Last commit message
Last commit date
VoiceStudio ranking on TrendshiftVoiceStudio logo

VoiceStudio

Previously OmniVoice-Studio

Clone voices, dub video, dictate, and produce long-form audio on your own hardware.

16 TTS engines · 11 ASR engines · 646-language catalogue · macOS, Windows, Linux, and Docker

No account, API key, subscription, or usage meter for the local workflow.

Install · Features · Compare · Requirements · Engines · Architecture · API · Docs · 简体中文

GitHub starsTotal downloadsLatest releaseAGPL-3.0 licenseDiscord community

Download VoiceStudio

Switching TTS engines from the VoiceStudio status bar

Warning

Active beta. Use the latest release for stable work. main contains the newest fixes and may change between releases. Report problems through GitHub Issues.

At a glance

VoiceStudio
WorkflowsVoice cloning and design, video dubbing, dictation, stories, audiobooks, batch generation
Language catalogue646 TTS languages; actual coverage and quality depend on the selected engine
Engines16 TTS · 11 ASR · switch in Model Catalogue or with Ctrl/Cmd+E
PlatformsmacOS 13.3+ on Apple Silicon · Windows 10/11 x64 · Linux x86_64 with glibc 2.39+
ComputeCUDA · Apple Silicon MPS/MLX · ROCm on Linux · CPU · optional remote workers
InterfacesDesktop app · local REST/SSE/WebSocket API · OpenAI-compatible audio API · MCP Server
StorageVoices, projects, settings, and outputs stay on the machine by default
LicenseAGPL-3.0 application; downloaded models keep their upstream terms

Install

Download a package from the latest release, then follow the platform guide.

PlatformPackageGuide
macOS 13.3+Apple Silicon DMGInstall on macOS
Windows 10/11x64 MSI; choose the current-user build when listed to install without admin accessInstall on Windows
LinuxAppImage, x86_64 with glibc 2.39+Install on Linux
DockerCUDA, ROCm, CPU, and worker-only GPU profilesRun with Docker

First launch creates a managed Python environment and downloads the default model. Later launches reuse both.

Note

On macOS, first launch needs a one-time right-click, then Open approval. Intel Macs cannot run the local Python backend; use a remote backend instead.

First voice

  1. Launch VoiceStudio and open Voice Cloning.
  2. Add a clean voice sample. Three seconds works; 5 to 15 seconds usually gives a better prompt.
  3. Enter text, choose a language, then select Generate.

Run from source

Install the development prerequisites, then:

git clone https://github.com/debpalash/VoiceStudio.git
cd VoiceStudio
bun install
bun run desktop

Use bun run dev for the browser UI. See Contributing for services, tests, and platform packages.

If setup fails

Features

AreaIncluded
Voice CloningZero-shot synthesis from a short reference clip
Voice DesignCreate a voice from age, accent, pitch, style, and delivery instructions
Video DubbingTranscribe, translate, preserve speakers, synthesize, and export video
Stories and audiobooksMulti-voice scripts · EPUB/PDF import · chapter rendering · .m4b export
Dictation WidgetSystem-wide shortcut, live transcription, optional local-LLM cleanup
Vocal IsolationDemucs speech/background separation
Speaker DiarizationPyannote and WhisperX speaker assignment
Batch QueueQueue large sets of audio and video jobs with per-job progress
Model CatalogueInstall, remove, select, and route TTS, ASR, and LLM models
Remote Model DownloadsInstall models on enrolled remote workers with live progress
GPU Auto-DetectCUDA, MPS, ROCm, and CPU routing with per-engine checks
AI WatermarkAudioSeal embedding and detection
MCP ServerSynthesis and transcription tools for MCP clients
DiagnosticsSelf-checks, error journal, logs, and scrubbed support bundles
Local-firstCore creation stays local; network-backed features are explicit opt-ins
ExtensibleRegistry-based TTS, ASR, and plugin interfaces
VoiceStudio Model CatalogueSaving a gallery voice as a local profile
Model Catalogue: engine, device, and install stateGallery: save a shared voice as a local profile

Comparison

VoiceStudio trades managed cloud compute for local control. This is the practical difference:

VoiceStudioTypical hosted voice service
Best fitPrivate, offline, self-hosted, or high-volume workFast setup without local model management
Data pathLocal by default; remote features are opt-inAudio and text are processed by the provider
Cost modelFree software; you supply the hardwareSubscription, credits, or metered API use
SetupInstall the app and model weightsCreate an account and use the web app or API
PerformanceDepends on your engine and hardwareProvider manages compute and scaling
Offline useYes, after required models are installedUsually requires a network connection
CustomizationSource, engines, models, API, and routing are openLimited to provider options
MaintenanceYou manage updates, disk, and computeProvider manages infrastructure

Requirements

Requirements vary by engine. These values cover the default local workflow.

MinimumRecommended
OSWindows 10 x64 · macOS 13.3 Apple Silicon · Linux x86_64 with glibc 2.39+Current supported OS release
RAM8 GB16 GB+
Disk10 GB free20 GB+ SSD
GPUOptional; CPU mode is supportedNVIDIA CUDA or Apple Silicon
VRAM4 GB when using a GPU8 GB+; large optional engines need more
Python from source3.11+3.11 or 3.12

ROCm is Linux-only and opt-in. Windows AMD/Ryzen AI uses CPU. Systems with limited VRAM offload work to CPU when required. See performance, benchmarks, and engine disk usage.

Engines

Engine support is capability-specific. Check cloning, language, platform, memory, and license before choosing one. Full setup guides: docs/engines.

Text to speech

EngineLanguagesCloneInstructLinuxmacOS ARMWindowsLicense
VoiceStudio (default, powered by k2-fsa/OmniVoice)600+YesYesCUDA/CPUMPSCUDA/CPUAGPL-3.0 app · Apache-2.0 code, CC-BY-NC weights³
CosyVoice 39 + 18 dialectsYesYesCUDA/CPUCPUCUDA/CPUApache-2.0
GPT-SoVITS5YesNoCUDA/CPUNoCUDA/CPUMIT
VoxCPM230YesYesCUDA/CPUMPSCUDA/CPUApache-2.0
MOSS-TTS-Nano20YesNoCUDA/CPUCPUCUDA/CPUApache-2.0
KittenTTSEnglishNoNoCPUCPUCPUMIT
MLX-AudioModel-dependentVariesVariesNoMLXNoVaries
Sherpa-ONNX20+NoNoCUDA/CPUCPUCUDA/CPUApache-2.0
IndexTTS 2.5ZH · EN · JA · ES · ARYesNoCUDA/CPUCPUCUDA/CPUBilibili model license¹
OmniVoice GGUF600+YesYesCUDA/CPUMPS/CPUCUDA/CPUAGPL-3.0 app · review the derivative model terms³
OmniVoice (subprocess)600+YesYesCUDA/CPUMPSCUDA/CPUAGPL-3.0 app · Apache-2.0 code, CC-BY-NC weights³
PocketTTSEN · FR · DE · PT · IT · ESYesNoCPUCPUCPUCC-BY-4.0, gated²
Supertonic 331NoNoCPUCPUCPUOpenRAIL-M
MOSS-TTS-v1.531YesNoCUDA/CPUCPUCUDA/CPUApache-2.0
dots.tts24YesNoCUDA/CPUCPUNoApache-2.0
Confucius4-TTS14YesNoCUDA/CPUCPUCUDA/CPUApache-2.0

⚡ Installed or registered on demand.

¹ IndexTTS 2.5 requires a separate written Bilibili license above 100 million monthly active users or RMB 1 billion annual revenue. Review the model license.

² PocketTTS shows its gated-access and CC-BY-4.0 terms before first use.

³ The OmniVoice snapshot also includes an audio tokenizer under separate Boson Higgs Audio 2 and Meta Llama community terms. VoiceStudio's application license does not replace model or tokenizer terms.

Clone-less engines cannot preserve a reference speaker in dubbing or pinned-voice batch jobs. VoiceStudio rejects those jobs instead of silently changing engines. Heavy engines have separate memory and platform limits; check their engine guide first.

Speech to text

EngineIDLanguagesBest fit
WhisperX (default)whisperx~100Dubbing, subtitles, word-level timing
Faster-Whisperfaster-whisper~100General cross-platform transcription
Faster-Whisper (isolated)faster-whisper-isolated~100Crash-isolated batch transcription
MLX Whispermlx-whisper~100Apple Silicon
PyTorch Whisperpytorch-whisper~100CUDA, MPS, and CPU fallback
Parakeet TDTnemo-parakeetEnglish + 25 EUFast CPU/CUDA transcription
Parakeet TDT v3 (MLX)parakeet-mlx25 EUApple Silicon dictation and word timestamps
MoonshinemoonshineEnglishLow-power, low-latency ONNX
FunASRfunasr50+VAD and inline diarization
sherpa-onnx (live dictation)sherpa-onnx-asrModel-dependentStreaming CPU dictation
OpenAI-compatible⚠️ configured serveropenai-compat-asrServer-dependentLocal gigastt/Qwen3-ASR or a remote endpoint; audio goes only to that server

WhisperX and Faster-Whisper retry with int8 when efficient float16 is unavailable. Pin ASR_COMPUTE_TYPE=int8 or float32 only if automatic selection still fails.

Architecture

Tauri v2 desktop shell (Rust)
│ IPC
React + Vite UI
│ HTTP · SSE · WebSocket on localhost:3900
FastAPI backend
├── TTS / ASR engine registries
├── dubbing / audio / long-form pipelines
├── OpenAI-compatible API and MCP server
└── SQLite + Alembic → omnivoice_data/
LayerPathResponsibility
Desktop shellfrontend/src-tauri/Window lifecycle, tray, shortcuts, updater, sidecar bootstrap
Frontendfrontend/src/React UI, Zustand state, API and event clients, i18n
APIbackend/api/REST routes, schemas, auth boundaries, streaming
Core servicesbackend/services/Generation, dubbing, audio processing, persistence
Enginesbackend/engines/Isolated and optional engine adapters
Worker systembackend/worker/Authenticated remote compute and job transport
Dataomnivoice_data/Projects, voices, settings, logs, and SQLite state
Deliveryscripts/, deploy/, .github/workflows/Development, packaging, containers, releases, CI

Network boundary

  • The desktop talks to a loopback-only backend on localhost:3900.
  • Loopback API calls need no server key. Remote access requires a share PIN or API key.
  • Remote workers and OpenAI-compatible ASR are opt-in. Loopback ASR may use HTTP and keeps audio on the machine; non-loopback endpoints require HTTPS, and redirects are not followed.
  • Analytics is off until consent. If enabled, it sends allowlisted, content-free usage metadata. It never sends text, audio, file names, or projects.

Local speech platform and OpenAI-compatible API

Point an OpenAI-compatible audio client at the local backend:

- base_url="https://api.openai.com/v1"+ base_url="http://localhost:3900/v1"
EndpointPurpose
POST /v1/audio/speechTTS to mp3, opus, aac, flac, wav, or pcm; select a profile with voice and an engine with model
POST /v1/audio/transcriptionsSTT to json, text, verbose_json, srt, or vtt
WS /v1/audio/transcriptions/streamLive PCM/WebM transcription with partial, utterance, and session-final events
GET /.well-known/voicestudio-speechDiscover HTTP, WebSocket, MCP, and native dictation-control transports
GET /v1/audio/voicesList local voice profiles and engines
fromopenaiimportOpenAIclient=OpenAI(base_url="http://localhost:3900/v1", api_key="local")
withclient.audio.speech.with_streaming_response.create(
model="tts-1",
voice="<profile-id>",
input="Made on my own hardware.",
response_format="wav",
) asresponse:
response.stream_to_file("speech.wav")

The bundled Rust control sidecar lets Herdr, coding agents, VS Code, desktop apps, and TUIs trigger the system-wide dictation flow or reuse its native text insertion. See the speech platform guide. The full API reference is in Settings → OpenAPI Reference. For LAN, Tailscale, or proxy access, read API authentication before exposing the backend.

Agent skills

Install the VoiceStudio skills for Claude Code, Codex, Cursor, and other skills.sh-compatible agents:

npx skills add debpalash/VoiceStudio
  • omnivoice: synthesize speech and transcribe audio through local VoiceStudio.
  • oss-maintainer: the repository's open-source maintenance workflow.

Google Colab

Open in Colab

The notebook runs the app and web UI on a Colab GPU. Colab is remote compute, so uploaded audio and project data do not remain local to your machine.

Documentation

NeedRead
InstallmacOS · Windows · Linux · Docker
Fix setupTroubleshooting · model downloads · Hugging Face token
Choose an engineEngine guides · benchmarks · expressive speech
Tune hardwarePerformance · remote workers
Build integrationsSpeech platform · Private production API · API auth · MCP · examples
Build VoiceStudioContributing · engine acceptance
Track changesChangelog · roadmap · latest release
Remove everythingUninstall guide

FAQ

Does it work on Apple Silicon and Intel Macs?

Apple Silicon is supported with MPS and MLX options. Intel Macs cannot run the local backend because current PyTorch wheels are unavailable; they can connect to a remote backend. See macOS installation.

How much VRAM do I need?

A GPU is optional. Use 4 GB VRAM as the minimum for accelerated work and 8 GB+ for the default multi-stage workflow. Large optional engines can require 12 to 16 GB or more. Check the benchmarks and engine guide.

Why does a longer reference clip not always improve the clone?

Cloning is zero-shot: the clip is a prompt, not training data. Use 5 to 15 seconds of one speaker, close to the microphone, without music, noise, or reverb. Match the tone and pace you want in the output. For training, see data preparation and training.

Can I use generated audio commercially?

VoiceStudio's application license does not restrict generated audio, but it does not grant rights under a model's separate terms. The default OmniVoice repository labels its pretrained weights CC-BY-NC and includes a tokenizer under separate community terms. Review the selected model terms before commercial use.

Does VoiceStudio collect data?

Not unless you opt in. Analytics is off by default and skipping consent keeps it off. When enabled, the app sends allowlisted, content-free usage metadata. Text, audio, file names, voices, and projects are excluded. Change this at Settings → Privacy.

How do I remove VoiceStudio and its data?

Use scripts/uninstall.sh on macOS/Linux or scripts\uninstall.ps1 on Windows. Both show a dry run before deletion. See the uninstall guide for every path.

Community and contributing

Support development

VoiceStudio is free and has no paid tier. Donations fund development and infrastructure.

Ko-fi · PayPal · Sponsorship details

License

VoiceStudio is licensed under AGPL-3.0. You may run it, modify it, and use it internally. The application license itself does not restrict selling generated audio, but downloaded model and tokenizer terms may. If you modify VoiceStudio and provide that modified version as a network service, AGPL requires you to offer the corresponding source under the same license. A commercial license for VoiceStudio-owned code is available for proprietary embedding; it does not relicense third-party models. Contact VoiceStudio@palash.dev. See LICENSE-NOTICE.md for the plain-language scope.

Optional engines and downloaded models retain their own licenses. The bundled omnivoice/ Python code is Apache-2.0 upstream; the default downloaded weights and audio tokenizer use separate terms.

Acknowledgments

VoiceStudio builds on OmniVoice, WhisperX, Demucs, Pyannote, CTranslate2, AudioSeal, Tauri, Supertonic, Sherpa-ONNX, GPT-SoVITS, and PocketTTS.

About

VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

14.0k stars

Watchers

67 watching

Forks

Releases

Sponsor this project

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

1,893 Commits

Folders and files

NameName
Last commit message
Last commit date
VoiceStudio ranking on TrendshiftVoiceStudio logo

VoiceStudio

Previously OmniVoice-Studio

Clone voices, dub video, dictate, and produce long-form audio on your own hardware.

16 TTS engines · 11 ASR engines · 646-language catalogue · macOS, Windows, Linux, and Docker

No account, API key, subscription, or usage meter for the local workflow.

Install · Features · Compare · Requirements · Engines · Architecture · API · Docs · 简体中文

GitHub starsTotal downloadsLatest releaseAGPL-3.0 licenseDiscord community

Download VoiceStudio

Switching TTS engines from the VoiceStudio status bar

Warning

Active beta. Use the latest release for stable work. main contains the newest fixes and may change between releases. Report problems through GitHub Issues.

At a glance

VoiceStudio
WorkflowsVoice cloning and design, video dubbing, dictation, stories, audiobooks, batch generation
Language catalogue646 TTS languages; actual coverage and quality depend on the selected engine
Engines16 TTS · 11 ASR · switch in Model Catalogue or with Ctrl/Cmd+E
PlatformsmacOS 13.3+ on Apple Silicon · Windows 10/11 x64 · Linux x86_64 with glibc 2.39+
ComputeCUDA · Apple Silicon MPS/MLX · ROCm on Linux · CPU · optional remote workers
InterfacesDesktop app · local REST/SSE/WebSocket API · OpenAI-compatible audio API · MCP Server
StorageVoices, projects, settings, and outputs stay on the machine by default
LicenseAGPL-3.0 application; downloaded models keep their upstream terms

Install

Download a package from the latest release, then follow the platform guide.

PlatformPackageGuide
macOS 13.3+Apple Silicon DMGInstall on macOS
Windows 10/11x64 MSI; choose the current-user build when listed to install without admin accessInstall on Windows
LinuxAppImage, x86_64 with glibc 2.39+Install on Linux
DockerCUDA, ROCm, CPU, and worker-only GPU profilesRun with Docker

First launch creates a managed Python environment and downloads the default model. Later launches reuse both.

Note

On macOS, first launch needs a one-time right-click, then Open approval. Intel Macs cannot run the local Python backend; use a remote backend instead.

First voice

  1. Launch VoiceStudio and open Voice Cloning.
  2. Add a clean voice sample. Three seconds works; 5 to 15 seconds usually gives a better prompt.
  3. Enter text, choose a language, then select Generate.

Run from source

Install the development prerequisites, then:

git clone https://github.com/debpalash/VoiceStudio.git
cd VoiceStudio
bun install
bun run desktop

Use bun run dev for the browser UI. See Contributing for services, tests, and platform packages.

If setup fails

Features

AreaIncluded
Voice CloningZero-shot synthesis from a short reference clip
Voice DesignCreate a voice from age, accent, pitch, style, and delivery instructions
Video DubbingTranscribe, translate, preserve speakers, synthesize, and export video
Stories and audiobooksMulti-voice scripts · EPUB/PDF import · chapter rendering · .m4b export
Dictation WidgetSystem-wide shortcut, live transcription, optional local-LLM cleanup
Vocal IsolationDemucs speech/background separation
Speaker DiarizationPyannote and WhisperX speaker assignment
Batch QueueQueue large sets of audio and video jobs with per-job progress
Model CatalogueInstall, remove, select, and route TTS, ASR, and LLM models
Remote Model DownloadsInstall models on enrolled remote workers with live progress
GPU Auto-DetectCUDA, MPS, ROCm, and CPU routing with per-engine checks
AI WatermarkAudioSeal embedding and detection
MCP ServerSynthesis and transcription tools for MCP clients
DiagnosticsSelf-checks, error journal, logs, and scrubbed support bundles
Local-firstCore creation stays local; network-backed features are explicit opt-ins
ExtensibleRegistry-based TTS, ASR, and plugin interfaces
VoiceStudio Model CatalogueSaving a gallery voice as a local profile
Model Catalogue: engine, device, and install stateGallery: save a shared voice as a local profile

Comparison

VoiceStudio trades managed cloud compute for local control. This is the practical difference:

VoiceStudioTypical hosted voice service
Best fitPrivate, offline, self-hosted, or high-volume workFast setup without local model management
Data pathLocal by default; remote features are opt-inAudio and text are processed by the provider
Cost modelFree software; you supply the hardwareSubscription, credits, or metered API use
SetupInstall the app and model weightsCreate an account and use the web app or API
PerformanceDepends on your engine and hardwareProvider manages compute and scaling
Offline useYes, after required models are installedUsually requires a network connection
CustomizationSource, engines, models, API, and routing are openLimited to provider options
MaintenanceYou manage updates, disk, and computeProvider manages infrastructure

Requirements

Requirements vary by engine. These values cover the default local workflow.

MinimumRecommended
OSWindows 10 x64 · macOS 13.3 Apple Silicon · Linux x86_64 with glibc 2.39+Current supported OS release
RAM8 GB16 GB+
Disk10 GB free20 GB+ SSD
GPUOptional; CPU mode is supportedNVIDIA CUDA or Apple Silicon
VRAM4 GB when using a GPU8 GB+; large optional engines need more
Python from source3.11+3.11 or 3.12

ROCm is Linux-only and opt-in. Windows AMD/Ryzen AI uses CPU. Systems with limited VRAM offload work to CPU when required. See performance, benchmarks, and engine disk usage.

Engines

Engine support is capability-specific. Check cloning, language, platform, memory, and license before choosing one. Full setup guides: docs/engines.

Text to speech

EngineLanguagesCloneInstructLinuxmacOS ARMWindowsLicense
VoiceStudio (default, powered by k2-fsa/OmniVoice)600+YesYesCUDA/CPUMPSCUDA/CPUAGPL-3.0 app · Apache-2.0 code, CC-BY-NC weights³
CosyVoice 39 + 18 dialectsYesYesCUDA/CPUCPUCUDA/CPUApache-2.0
GPT-SoVITS5YesNoCUDA/CPUNoCUDA/CPUMIT
VoxCPM230YesYesCUDA/CPUMPSCUDA/CPUApache-2.0
MOSS-TTS-Nano20YesNoCUDA/CPUCPUCUDA/CPUApache-2.0
KittenTTSEnglishNoNoCPUCPUCPUMIT
MLX-AudioModel-dependentVariesVariesNoMLXNoVaries
Sherpa-ONNX20+NoNoCUDA/CPUCPUCUDA/CPUApache-2.0
IndexTTS 2.5ZH · EN · JA · ES · ARYesNoCUDA/CPUCPUCUDA/CPUBilibili model license¹
OmniVoice GGUF600+YesYesCUDA/CPUMPS/CPUCUDA/CPUAGPL-3.0 app · review the derivative model terms³
OmniVoice (subprocess)600+YesYesCUDA/CPUMPSCUDA/CPUAGPL-3.0 app · Apache-2.0 code, CC-BY-NC weights³
PocketTTSEN · FR · DE · PT · IT · ESYesNoCPUCPUCPUCC-BY-4.0, gated²
Supertonic 331NoNoCPUCPUCPUOpenRAIL-M
MOSS-TTS-v1.531YesNoCUDA/CPUCPUCUDA/CPUApache-2.0
dots.tts24YesNoCUDA/CPUCPUNoApache-2.0
Confucius4-TTS14YesNoCUDA/CPUCPUCUDA/CPUApache-2.0

⚡ Installed or registered on demand.

¹ IndexTTS 2.5 requires a separate written Bilibili license above 100 million monthly active users or RMB 1 billion annual revenue. Review the model license.

² PocketTTS shows its gated-access and CC-BY-4.0 terms before first use.

³ The OmniVoice snapshot also includes an audio tokenizer under separate Boson Higgs Audio 2 and Meta Llama community terms. VoiceStudio's application license does not replace model or tokenizer terms.

Clone-less engines cannot preserve a reference speaker in dubbing or pinned-voice batch jobs. VoiceStudio rejects those jobs instead of silently changing engines. Heavy engines have separate memory and platform limits; check their engine guide first.

Speech to text

EngineIDLanguagesBest fit
WhisperX (default)whisperx~100Dubbing, subtitles, word-level timing
Faster-Whisperfaster-whisper~100General cross-platform transcription
Faster-Whisper (isolated)faster-whisper-isolated~100Crash-isolated batch transcription
MLX Whispermlx-whisper~100Apple Silicon
PyTorch Whisperpytorch-whisper~100CUDA, MPS, and CPU fallback
Parakeet TDTnemo-parakeetEnglish + 25 EUFast CPU/CUDA transcription
Parakeet TDT v3 (MLX)parakeet-mlx25 EUApple Silicon dictation and word timestamps
MoonshinemoonshineEnglishLow-power, low-latency ONNX
FunASRfunasr50+VAD and inline diarization
sherpa-onnx (live dictation)sherpa-onnx-asrModel-dependentStreaming CPU dictation
OpenAI-compatible⚠️ configured serveropenai-compat-asrServer-dependentLocal gigastt/Qwen3-ASR or a remote endpoint; audio goes only to that server

WhisperX and Faster-Whisper retry with int8 when efficient float16 is unavailable. Pin ASR_COMPUTE_TYPE=int8 or float32 only if automatic selection still fails.

Architecture

Tauri v2 desktop shell (Rust)
│ IPC
React + Vite UI
│ HTTP · SSE · WebSocket on localhost:3900
FastAPI backend
├── TTS / ASR engine registries
├── dubbing / audio / long-form pipelines
├── OpenAI-compatible API and MCP server
└── SQLite + Alembic → omnivoice_data/
LayerPathResponsibility
Desktop shellfrontend/src-tauri/Window lifecycle, tray, shortcuts, updater, sidecar bootstrap
Frontendfrontend/src/React UI, Zustand state, API and event clients, i18n
APIbackend/api/REST routes, schemas, auth boundaries, streaming
Core servicesbackend/services/Generation, dubbing, audio processing, persistence
Enginesbackend/engines/Isolated and optional engine adapters
Worker systembackend/worker/Authenticated remote compute and job transport
Dataomnivoice_data/Projects, voices, settings, logs, and SQLite state
Deliveryscripts/, deploy/, .github/workflows/Development, packaging, containers, releases, CI

Network boundary

  • The desktop talks to a loopback-only backend on localhost:3900.
  • Loopback API calls need no server key. Remote access requires a share PIN or API key.
  • Remote workers and OpenAI-compatible ASR are opt-in. Loopback ASR may use HTTP and keeps audio on the machine; non-loopback endpoints require HTTPS, and redirects are not followed.
  • Analytics is off until consent. If enabled, it sends allowlisted, content-free usage metadata. It never sends text, audio, file names, or projects.

Local speech platform and OpenAI-compatible API

Point an OpenAI-compatible audio client at the local backend:

- base_url="https://api.openai.com/v1"+ base_url="http://localhost:3900/v1"
EndpointPurpose
POST /v1/audio/speechTTS to mp3, opus, aac, flac, wav, or pcm; select a profile with voice and an engine with model
POST /v1/audio/transcriptionsSTT to json, text, verbose_json, srt, or vtt
WS /v1/audio/transcriptions/streamLive PCM/WebM transcription with partial, utterance, and session-final events
GET /.well-known/voicestudio-speechDiscover HTTP, WebSocket, MCP, and native dictation-control transports
GET /v1/audio/voicesList local voice profiles and engines
fromopenaiimportOpenAIclient=OpenAI(base_url="http://localhost:3900/v1", api_key="local")
withclient.audio.speech.with_streaming_response.create(
model="tts-1",
voice="<profile-id>",
input="Made on my own hardware.",
response_format="wav",
) asresponse:
response.stream_to_file("speech.wav")

The bundled Rust control sidecar lets Herdr, coding agents, VS Code, desktop apps, and TUIs trigger the system-wide dictation flow or reuse its native text insertion. See the speech platform guide. The full API reference is in Settings → OpenAPI Reference. For LAN, Tailscale, or proxy access, read API authentication before exposing the backend.

Agent skills

Install the VoiceStudio skills for Claude Code, Codex, Cursor, and other skills.sh-compatible agents:

npx skills add debpalash/VoiceStudio
  • omnivoice: synthesize speech and transcribe audio through local VoiceStudio.
  • oss-maintainer: the repository's open-source maintenance workflow.

Google Colab

Open in Colab

The notebook runs the app and web UI on a Colab GPU. Colab is remote compute, so uploaded audio and project data do not remain local to your machine.

Documentation

NeedRead
InstallmacOS · Windows · Linux · Docker
Fix setupTroubleshooting · model downloads · Hugging Face token
Choose an engineEngine guides · benchmarks · expressive speech
Tune hardwarePerformance · remote workers
Build integrationsSpeech platform · Private production API · API auth · MCP · examples
Build VoiceStudioContributing · engine acceptance
Track changesChangelog · roadmap · latest release
Remove everythingUninstall guide

FAQ

Does it work on Apple Silicon and Intel Macs?

Apple Silicon is supported with MPS and MLX options. Intel Macs cannot run the local backend because current PyTorch wheels are unavailable; they can connect to a remote backend. See macOS installation.

How much VRAM do I need?

A GPU is optional. Use 4 GB VRAM as the minimum for accelerated work and 8 GB+ for the default multi-stage workflow. Large optional engines can require 12 to 16 GB or more. Check the benchmarks and engine guide.

Why does a longer reference clip not always improve the clone?

Cloning is zero-shot: the clip is a prompt, not training data. Use 5 to 15 seconds of one speaker, close to the microphone, without music, noise, or reverb. Match the tone and pace you want in the output. For training, see data preparation and training.

Can I use generated audio commercially?

VoiceStudio's application license does not restrict generated audio, but it does not grant rights under a model's separate terms. The default OmniVoice repository labels its pretrained weights CC-BY-NC and includes a tokenizer under separate community terms. Review the selected model terms before commercial use.

Does VoiceStudio collect data?

Not unless you opt in. Analytics is off by default and skipping consent keeps it off. When enabled, the app sends allowlisted, content-free usage metadata. Text, audio, file names, voices, and projects are excluded. Change this at Settings → Privacy.

How do I remove VoiceStudio and its data?

Use scripts/uninstall.sh on macOS/Linux or scripts\uninstall.ps1 on Windows. Both show a dry run before deletion. See the uninstall guide for every path.

Community and contributing

Support development

VoiceStudio is free and has no paid tier. Donations fund development and infrastructure.

Ko-fi · PayPal · Sponsorship details

License

VoiceStudio is licensed under AGPL-3.0. You may run it, modify it, and use it internally. The application license itself does not restrict selling generated audio, but downloaded model and tokenizer terms may. If you modify VoiceStudio and provide that modified version as a network service, AGPL requires you to offer the corresponding source under the same license. A commercial license for VoiceStudio-owned code is available for proprietary embedding; it does not relicense third-party models. Contact VoiceStudio@palash.dev. See LICENSE-NOTICE.md for the plain-language scope.

Optional engines and downloaded models retain their own licenses. The bundled omnivoice/ Python code is Apache-2.0 upstream; the default downloaded weights and audio tokenizer use separate terms.

Acknowledgments

VoiceStudio builds on OmniVoice, WhisperX, Demucs, Pyannote, CTranslate2, AudioSeal, Tauri, Supertonic, Sherpa-ONNX, GPT-SoVITS, and PocketTTS.

About

VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

14.0k stars

Watchers

67 watching

Forks

Releases

Sponsor this project

Packages

Used by

Contributors

Languages