An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
-
Updated
Aug 18, 2026 - Python
An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
Audiobook creation app supporting too many TTS models (Qwen3-TTS, OmniVoice, VibeVoice, etc), focused on high-quality output. Plus audio-synced reader web app and standalone server component.
MLX implementation of IndexTTS (v1.5 & v2.0) — high-quality text-to-speech with voice cloning and emotion control, optimized for Apple Silicon
多引擎语音合成平台:VoxCPM2 与 IndexTTS2 引擎,声音克隆、音色设计、LoRA 微调、多角色剧本配音、情感控制,OpenAI 兼容 API,中英日韩界面,一键安装启动 | Multi-engine TTS platform: VoxCPM2 + IndexTTS2, voice cloning, voice design, LoRA fine-tuning, multi-character dubbing, emotion control, OpenAI-compatible API, zh/en/ja/ko UI, one-click install
Windows launcher and unified workbench for IndexTTS 2.5 voice cloning and TTS.
IndexTTS2 精确停顿控制([pause:N])ComfyUI 节点包:句中与句号前后停顿精确调整,平均偏差 13ms;支持批量生成、候选试听验收、SRT 字幕与 seed 复现。Precise pause control ([pause:N]) ComfyUI nodes for IndexTTS2 — waveform-domain pause editing in-sentence and around periods, ±13ms, with batch generation, candidate acceptance and SRT subtitles.
T8star-Aix IndexTTS 2.5 Windows Electron desktop integration
IndexTTS 2.5 MLX 8-bit for Apple Silicon with an official /gen_single Gradio-compatible service on port 7862
Precise pause control ([pause:Nms]) for IndexTTS 2.5: whisper dual-anchor localization + energy valley insertion + optional code-level pause detector (LR+LSTM). Minimal 7-line patch for official webui.py. 中英双语。
Bilingual IndexTTS 1.5 API service with local WebUI, voice management, and queued generation / IndexTTS 1.5 双语 API 服务,集成本地 WebUI、音色管理与队列式生成
AI Agent 可复用的本地语音生成技能 —— 通过 gradio_client 调用本机正在运行的 IndexTTS2 WebUI,复用已加载模型生成口播/配音(不重复加载、不占双份显存)。
Open a GitHub issue, get zero-shot voice-cloned speech back as a Release — IndexTTS-2 running on a free Google Colab T4 GPU, orchestrated entirely by GitHub Actions.
One-click IndexTTS deployment for RunPod
Production-derived TTS and HLS generation engine for long-form AI dubbing.
Archived: superseded by ComfyUI-Hydra-InferWorks v1.0.0
An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
Local multi-voice audiobook pipeline for the Shadow Slave web novel - LLM diarization + IndexTTS2 emotional voice cloning on a single 12 GB GPU
Run high-quality text-to-speech models locally on your CPU with low latency and no GPU requirements.
Add a description, image, and links to the indextts topic page so that developers can more easily learn about it.
To associate your repository with the indextts topic, visit your repo's landing page and select "manage topics."