Skip to content

Latest commit

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

AutoDub - Automatic Video Translation & Dubbing

Automatically transcribe, translate, and dub videos into different languages using AI-powered text-to-speech.

Features

  • 🎙️ Speech Recognition: Transcribe audio using OpenAI Whisper
  • 🌍 Translation: Translate to 100+ languages via Google Translate
  • 🗣️ Three TTS Engines:
    • Edge TTS: High-quality Microsoft voices (recommended)
    • Silero: Fast Russian TTS (offline after first download)
    • XTTS: Voice cloning from 6-10 second samples
  • 🎬 Video Preservation: Keeps original video, mixes original audio (20%) with dubbed audio (150%)
  • 📝 Subtitle Generation: Creates SRT files for translated text

Requirements

System Dependencies

# Fedora/RHEL
sudo dnf install ffmpeg python3.10 python3.10-devel
# Ubuntu/Debian
sudo apt install ffmpeg python3.10 python3.10-devel
# macOS
brew install ffmpeg

Python Dependencies

python3 -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
pip install torch torchaudio --index-url https://download.pytorch.org/whl/cpu
pip install openai-whisper pysrt edge-tts deep-translator soundfile tqdm
pip install TTS # Only needed for XTTS voice cloning

Local Translation with Ollama (Optional)

AutoDub now supports fully offline translation using Ollama. This is ideal for privacy, avoiding API limits, and achieving more context-aware translations.

1. Install Ollama

For Linux (Fedora/Ubuntu/etc.):

curl -fsSL [https://ollama.com/install.sh](https://ollama.com/install.sh) | sh
ollama pull llama3

Quick Start

Here is the concise guide on how to get started using your setup.sh script, formatted in Markdown: 🚀 Quick Start Guide

Follow these three steps to set up and start dubbing your videos:

  1. Prepare Files

Ensure you have the following files in your project directory:

setup.sh (The installer)
autodub_v4_1.py (The main engine)
install.txt (List of dependencies)
  1. Run Installation

Open your terminal in the project folder and execute:

chmod +x setup.sh && ./setup.sh

Basic Usage (Edge TTS - Recommended)

# Dub to Russian (default)
python autodub.py video.mp4
# Dub to English
python autodub.py video.mp4 --target_lang en
# Dub to German
python autodub.py video.mp4 --target_lang de

Silero TTS (Faster, Russian only)

# Default voice (aidar)
python autodub.py video.mp4 --tts silero
# Female voice
python autodub.py video.mp4 --tts silero --silero_voice xenia
# Available voices: aidar, baya, kseniya, xenia, eugene

XTTS Voice Cloning (Most Natural)

# Requires 6-10 second clean voice sample
python autodub.py video.mp4 --tts xtts --ref_voice my_voice.wav --target_lang en

Ollama translator

# Use Ollama with default llama3 model
./run.sh video.mp4 --translator ollama
# Use a specific model (e.g., Mistral)
./run.sh video.mp4 --translator ollama --ollama_model mistral 

Command-Line Options

positional arguments:
video Input video file
options:
-h, --help Show help message
--tts {edge,silero,xtts}
TTS engine (default: edge)
--target_lang LANG Target language code (default: ru)
Supports: ru, en, de, fr, es, it, pt, ja, zh, etc.
--silero_voice {aidar,baya,kseniya,xenia,eugene}
Silero voice for Russian (default: aidar)
--ref_voice FILE Reference WAV for XTTS voice cloning
--keep-temp Keep temporary files after processing

Supported Languages

Edge TTS supports 100+ languages. Common codes:

  • ru - Russian
  • en - English
  • de - German
  • fr - French
  • es - Spanish
  • it - Italian
  • pt - Portuguese
  • ja - Japanese
  • zh - Chinese

Full list: https://speech.microsoft.com/portal/voicegallery

Output

The script generates:

  • {video}_dubbed.mp4 - Video with dubbed audio
  • {video}_{lang}.srt - Subtitle file with translations

Performance

EngineSpeedQualityLanguagesNotes
Edge TTSFast⭐⭐⭐⭐⭐100+Best quality, requires internet
SileroVery Fast⭐⭐⭐Russian onlyOffline, robotic
XTTSSlow⭐⭐⭐⭐⭐16Voice cloning, GPU recommended

Troubleshooting

"No module named 'soundfile'"

pip install soundfile

"TorchCodec is required"

This is already patched in the code. If you still see it, update PyTorch:

pip install --upgrade torch torchaudio

Silero model download fails

The script will auto-download on first run (~40MB). Check your internet connection.

XTTS out of memory

Use CPU mode or reduce video length. For long videos, split into segments.

Poor voice quality with Silero

Use Edge TTS or XTTS instead. Silero is designed for speed, not quality.

Ollama Integration Features

  1. Privacy: Your transcripts and translations never leave your local machine.
  2. Custom Context: LLMs can handle nuances, slang, and technical terms better than basic translators.
  3. Cost: 100% free with no character limits or subscription fees.
  4. Offline Workflow: Combined with Silero or XTTS, you can dub videos without an active internet connection.
FeatureGoogle TranslateOllama (Local LLM)
SpeedInstantDepends on your GPU/RAM
SetupZero setupRequires model download
InternetRequiredNot required
QualityLiteral / StandardContextual / Natural

Technical Details

Processing Pipeline

  1. Extract Audio: FFmpeg extracts mono 16kHz WAV
  2. Transcribe: Whisper "base" model transcribes with timestamps
  3. Translate: Google Translate API translates segments
  4. Synthesize: TTS engine generates speech for each subtitle
  5. Merge: FFmpeg mixes original (20%) + dubbed (150%) audio with video

Audio Mixing

  • Original audio: 20% volume (background)
  • Dubbed audio: 150% volume (foreground)
  • Output: AAC 128kbps, video copied without re-encoding

License

MIT License - see LICENSE file

Credits

Contributing

Issues and pull requests welcome!

About

Automatic video translator and dubber using Whisper, XTTS v2 for voice cloning, and Ollama for local LLM translation. Supports 100+ languages.

Topics

Resources

Stars

24 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - shyhirt/AutoDub: Automatic video translator and dubber using Whisper, XTTS v2 for voice cloning, and Ollama for local LLM translation. Supports 100+ languages. · GitHub
Skip to content

Latest commit

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

AutoDub - Automatic Video Translation & Dubbing

Automatically transcribe, translate, and dub videos into different languages using AI-powered text-to-speech.

Features

  • 🎙️ Speech Recognition: Transcribe audio using OpenAI Whisper
  • 🌍 Translation: Translate to 100+ languages via Google Translate
  • 🗣️ Three TTS Engines:
    • Edge TTS: High-quality Microsoft voices (recommended)
    • Silero: Fast Russian TTS (offline after first download)
    • XTTS: Voice cloning from 6-10 second samples
  • 🎬 Video Preservation: Keeps original video, mixes original audio (20%) with dubbed audio (150%)
  • 📝 Subtitle Generation: Creates SRT files for translated text

Requirements

System Dependencies

# Fedora/RHEL
sudo dnf install ffmpeg python3.10 python3.10-devel
# Ubuntu/Debian
sudo apt install ffmpeg python3.10 python3.10-devel
# macOS
brew install ffmpeg

Python Dependencies

python3 -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
pip install torch torchaudio --index-url https://download.pytorch.org/whl/cpu
pip install openai-whisper pysrt edge-tts deep-translator soundfile tqdm
pip install TTS # Only needed for XTTS voice cloning

Local Translation with Ollama (Optional)

AutoDub now supports fully offline translation using Ollama. This is ideal for privacy, avoiding API limits, and achieving more context-aware translations.

1. Install Ollama

For Linux (Fedora/Ubuntu/etc.):

curl -fsSL [https://ollama.com/install.sh](https://ollama.com/install.sh) | sh
ollama pull llama3

Quick Start

Here is the concise guide on how to get started using your setup.sh script, formatted in Markdown: 🚀 Quick Start Guide

Follow these three steps to set up and start dubbing your videos:

  1. Prepare Files

Ensure you have the following files in your project directory:

setup.sh (The installer)
autodub_v4_1.py (The main engine)
install.txt (List of dependencies)
  1. Run Installation

Open your terminal in the project folder and execute:

chmod +x setup.sh && ./setup.sh

Basic Usage (Edge TTS - Recommended)

# Dub to Russian (default)
python autodub.py video.mp4
# Dub to English
python autodub.py video.mp4 --target_lang en
# Dub to German
python autodub.py video.mp4 --target_lang de

Silero TTS (Faster, Russian only)

# Default voice (aidar)
python autodub.py video.mp4 --tts silero
# Female voice
python autodub.py video.mp4 --tts silero --silero_voice xenia
# Available voices: aidar, baya, kseniya, xenia, eugene

XTTS Voice Cloning (Most Natural)

# Requires 6-10 second clean voice sample
python autodub.py video.mp4 --tts xtts --ref_voice my_voice.wav --target_lang en

Ollama translator

# Use Ollama with default llama3 model
./run.sh video.mp4 --translator ollama
# Use a specific model (e.g., Mistral)
./run.sh video.mp4 --translator ollama --ollama_model mistral 

Command-Line Options

positional arguments:
video Input video file
options:
-h, --help Show help message
--tts {edge,silero,xtts}
TTS engine (default: edge)
--target_lang LANG Target language code (default: ru)
Supports: ru, en, de, fr, es, it, pt, ja, zh, etc.
--silero_voice {aidar,baya,kseniya,xenia,eugene}
Silero voice for Russian (default: aidar)
--ref_voice FILE Reference WAV for XTTS voice cloning
--keep-temp Keep temporary files after processing

Supported Languages

Edge TTS supports 100+ languages. Common codes:

  • ru - Russian
  • en - English
  • de - German
  • fr - French
  • es - Spanish
  • it - Italian
  • pt - Portuguese
  • ja - Japanese
  • zh - Chinese

Full list: https://speech.microsoft.com/portal/voicegallery

Output

The script generates:

  • {video}_dubbed.mp4 - Video with dubbed audio
  • {video}_{lang}.srt - Subtitle file with translations

Performance

EngineSpeedQualityLanguagesNotes
Edge TTSFast⭐⭐⭐⭐⭐100+Best quality, requires internet
SileroVery Fast⭐⭐⭐Russian onlyOffline, robotic
XTTSSlow⭐⭐⭐⭐⭐16Voice cloning, GPU recommended

Troubleshooting

"No module named 'soundfile'"

pip install soundfile

"TorchCodec is required"

This is already patched in the code. If you still see it, update PyTorch:

pip install --upgrade torch torchaudio

Silero model download fails

The script will auto-download on first run (~40MB). Check your internet connection.

XTTS out of memory

Use CPU mode or reduce video length. For long videos, split into segments.

Poor voice quality with Silero

Use Edge TTS or XTTS instead. Silero is designed for speed, not quality.

Ollama Integration Features

  1. Privacy: Your transcripts and translations never leave your local machine.
  2. Custom Context: LLMs can handle nuances, slang, and technical terms better than basic translators.
  3. Cost: 100% free with no character limits or subscription fees.
  4. Offline Workflow: Combined with Silero or XTTS, you can dub videos without an active internet connection.
FeatureGoogle TranslateOllama (Local LLM)
SpeedInstantDepends on your GPU/RAM
SetupZero setupRequires model download
InternetRequiredNot required
QualityLiteral / StandardContextual / Natural

Technical Details

Processing Pipeline

  1. Extract Audio: FFmpeg extracts mono 16kHz WAV
  2. Transcribe: Whisper "base" model transcribes with timestamps
  3. Translate: Google Translate API translates segments
  4. Synthesize: TTS engine generates speech for each subtitle
  5. Merge: FFmpeg mixes original (20%) + dubbed (150%) audio with video

Audio Mixing

  • Original audio: 20% volume (background)
  • Dubbed audio: 150% volume (foreground)
  • Output: AAC 128kbps, video copied without re-encoding

License

MIT License - see LICENSE file

Credits

Contributing

Issues and pull requests welcome!

About

Automatic video translator and dubber using Whisper, XTTS v2 for voice cloning, and Ollama for local LLM translation. Supports 100+ languages.

Topics

Resources

Stars

24 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - shyhirt/AutoDub: Automatic video translator and dubber using Whisper, XTTS v2 for voice cloning, and Ollama for local LLM translation. Supports 100+ languages. · GitHub
Skip to content

Latest commit

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

AutoDub - Automatic Video Translation & Dubbing

Automatically transcribe, translate, and dub videos into different languages using AI-powered text-to-speech.

Features

  • 🎙️ Speech Recognition: Transcribe audio using OpenAI Whisper
  • 🌍 Translation: Translate to 100+ languages via Google Translate
  • 🗣️ Three TTS Engines:
    • Edge TTS: High-quality Microsoft voices (recommended)
    • Silero: Fast Russian TTS (offline after first download)
    • XTTS: Voice cloning from 6-10 second samples
  • 🎬 Video Preservation: Keeps original video, mixes original audio (20%) with dubbed audio (150%)
  • 📝 Subtitle Generation: Creates SRT files for translated text

Requirements

System Dependencies

# Fedora/RHEL
sudo dnf install ffmpeg python3.10 python3.10-devel
# Ubuntu/Debian
sudo apt install ffmpeg python3.10 python3.10-devel
# macOS
brew install ffmpeg

Python Dependencies

python3 -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
pip install torch torchaudio --index-url https://download.pytorch.org/whl/cpu
pip install openai-whisper pysrt edge-tts deep-translator soundfile tqdm
pip install TTS # Only needed for XTTS voice cloning

Local Translation with Ollama (Optional)

AutoDub now supports fully offline translation using Ollama. This is ideal for privacy, avoiding API limits, and achieving more context-aware translations.

1. Install Ollama

For Linux (Fedora/Ubuntu/etc.):

curl -fsSL [https://ollama.com/install.sh](https://ollama.com/install.sh) | sh
ollama pull llama3

Quick Start

Here is the concise guide on how to get started using your setup.sh script, formatted in Markdown: 🚀 Quick Start Guide

Follow these three steps to set up and start dubbing your videos:

  1. Prepare Files

Ensure you have the following files in your project directory:

setup.sh (The installer)
autodub_v4_1.py (The main engine)
install.txt (List of dependencies)
  1. Run Installation

Open your terminal in the project folder and execute:

chmod +x setup.sh && ./setup.sh

Basic Usage (Edge TTS - Recommended)

# Dub to Russian (default)
python autodub.py video.mp4
# Dub to English
python autodub.py video.mp4 --target_lang en
# Dub to German
python autodub.py video.mp4 --target_lang de

Silero TTS (Faster, Russian only)

# Default voice (aidar)
python autodub.py video.mp4 --tts silero
# Female voice
python autodub.py video.mp4 --tts silero --silero_voice xenia
# Available voices: aidar, baya, kseniya, xenia, eugene

XTTS Voice Cloning (Most Natural)

# Requires 6-10 second clean voice sample
python autodub.py video.mp4 --tts xtts --ref_voice my_voice.wav --target_lang en

Ollama translator

# Use Ollama with default llama3 model
./run.sh video.mp4 --translator ollama
# Use a specific model (e.g., Mistral)
./run.sh video.mp4 --translator ollama --ollama_model mistral 

Command-Line Options

positional arguments:
video Input video file
options:
-h, --help Show help message
--tts {edge,silero,xtts}
TTS engine (default: edge)
--target_lang LANG Target language code (default: ru)
Supports: ru, en, de, fr, es, it, pt, ja, zh, etc.
--silero_voice {aidar,baya,kseniya,xenia,eugene}
Silero voice for Russian (default: aidar)
--ref_voice FILE Reference WAV for XTTS voice cloning
--keep-temp Keep temporary files after processing

Supported Languages

Edge TTS supports 100+ languages. Common codes:

  • ru - Russian
  • en - English
  • de - German
  • fr - French
  • es - Spanish
  • it - Italian
  • pt - Portuguese
  • ja - Japanese
  • zh - Chinese

Full list: https://speech.microsoft.com/portal/voicegallery

Output

The script generates:

  • {video}_dubbed.mp4 - Video with dubbed audio
  • {video}_{lang}.srt - Subtitle file with translations

Performance

EngineSpeedQualityLanguagesNotes
Edge TTSFast⭐⭐⭐⭐⭐100+Best quality, requires internet
SileroVery Fast⭐⭐⭐Russian onlyOffline, robotic
XTTSSlow⭐⭐⭐⭐⭐16Voice cloning, GPU recommended

Troubleshooting

"No module named 'soundfile'"

pip install soundfile

"TorchCodec is required"

This is already patched in the code. If you still see it, update PyTorch:

pip install --upgrade torch torchaudio

Silero model download fails

The script will auto-download on first run (~40MB). Check your internet connection.

XTTS out of memory

Use CPU mode or reduce video length. For long videos, split into segments.

Poor voice quality with Silero

Use Edge TTS or XTTS instead. Silero is designed for speed, not quality.

Ollama Integration Features

  1. Privacy: Your transcripts and translations never leave your local machine.
  2. Custom Context: LLMs can handle nuances, slang, and technical terms better than basic translators.
  3. Cost: 100% free with no character limits or subscription fees.
  4. Offline Workflow: Combined with Silero or XTTS, you can dub videos without an active internet connection.
FeatureGoogle TranslateOllama (Local LLM)
SpeedInstantDepends on your GPU/RAM
SetupZero setupRequires model download
InternetRequiredNot required
QualityLiteral / StandardContextual / Natural

Technical Details

Processing Pipeline

  1. Extract Audio: FFmpeg extracts mono 16kHz WAV
  2. Transcribe: Whisper "base" model transcribes with timestamps
  3. Translate: Google Translate API translates segments
  4. Synthesize: TTS engine generates speech for each subtitle
  5. Merge: FFmpeg mixes original (20%) + dubbed (150%) audio with video

Audio Mixing

  • Original audio: 20% volume (background)
  • Dubbed audio: 150% volume (foreground)
  • Output: AAC 128kbps, video copied without re-encoding

License

MIT License - see LICENSE file

Credits

Contributing

Issues and pull requests welcome!

About

Automatic video translator and dubber using Whisper, XTTS v2 for voice cloning, and Ollama for local LLM translation. Supports 100+ languages.

Topics

Resources

Stars

24 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - shyhirt/AutoDub: Automatic video translator and dubber using Whisper, XTTS v2 for voice cloning, and Ollama for local LLM translation. Supports 100+ languages. · GitHub
Skip to content

Latest commit

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

AutoDub - Automatic Video Translation & Dubbing

Automatically transcribe, translate, and dub videos into different languages using AI-powered text-to-speech.

Features

  • 🎙️ Speech Recognition: Transcribe audio using OpenAI Whisper
  • 🌍 Translation: Translate to 100+ languages via Google Translate
  • 🗣️ Three TTS Engines:
    • Edge TTS: High-quality Microsoft voices (recommended)
    • Silero: Fast Russian TTS (offline after first download)
    • XTTS: Voice cloning from 6-10 second samples
  • 🎬 Video Preservation: Keeps original video, mixes original audio (20%) with dubbed audio (150%)
  • 📝 Subtitle Generation: Creates SRT files for translated text

Requirements

System Dependencies

# Fedora/RHEL
sudo dnf install ffmpeg python3.10 python3.10-devel
# Ubuntu/Debian
sudo apt install ffmpeg python3.10 python3.10-devel
# macOS
brew install ffmpeg

Python Dependencies

python3 -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
pip install torch torchaudio --index-url https://download.pytorch.org/whl/cpu
pip install openai-whisper pysrt edge-tts deep-translator soundfile tqdm
pip install TTS # Only needed for XTTS voice cloning

Local Translation with Ollama (Optional)

AutoDub now supports fully offline translation using Ollama. This is ideal for privacy, avoiding API limits, and achieving more context-aware translations.

1. Install Ollama

For Linux (Fedora/Ubuntu/etc.):

curl -fsSL [https://ollama.com/install.sh](https://ollama.com/install.sh) | sh
ollama pull llama3

Quick Start

Here is the concise guide on how to get started using your setup.sh script, formatted in Markdown: 🚀 Quick Start Guide

Follow these three steps to set up and start dubbing your videos:

  1. Prepare Files

Ensure you have the following files in your project directory:

setup.sh (The installer)
autodub_v4_1.py (The main engine)
install.txt (List of dependencies)
  1. Run Installation

Open your terminal in the project folder and execute:

chmod +x setup.sh && ./setup.sh

Basic Usage (Edge TTS - Recommended)

# Dub to Russian (default)
python autodub.py video.mp4
# Dub to English
python autodub.py video.mp4 --target_lang en
# Dub to German
python autodub.py video.mp4 --target_lang de

Silero TTS (Faster, Russian only)

# Default voice (aidar)
python autodub.py video.mp4 --tts silero
# Female voice
python autodub.py video.mp4 --tts silero --silero_voice xenia
# Available voices: aidar, baya, kseniya, xenia, eugene

XTTS Voice Cloning (Most Natural)

# Requires 6-10 second clean voice sample
python autodub.py video.mp4 --tts xtts --ref_voice my_voice.wav --target_lang en

Ollama translator

# Use Ollama with default llama3 model
./run.sh video.mp4 --translator ollama
# Use a specific model (e.g., Mistral)
./run.sh video.mp4 --translator ollama --ollama_model mistral 

Command-Line Options

positional arguments:
video Input video file
options:
-h, --help Show help message
--tts {edge,silero,xtts}
TTS engine (default: edge)
--target_lang LANG Target language code (default: ru)
Supports: ru, en, de, fr, es, it, pt, ja, zh, etc.
--silero_voice {aidar,baya,kseniya,xenia,eugene}
Silero voice for Russian (default: aidar)
--ref_voice FILE Reference WAV for XTTS voice cloning
--keep-temp Keep temporary files after processing

Supported Languages

Edge TTS supports 100+ languages. Common codes:

  • ru - Russian
  • en - English
  • de - German
  • fr - French
  • es - Spanish
  • it - Italian
  • pt - Portuguese
  • ja - Japanese
  • zh - Chinese

Full list: https://speech.microsoft.com/portal/voicegallery

Output

The script generates:

  • {video}_dubbed.mp4 - Video with dubbed audio
  • {video}_{lang}.srt - Subtitle file with translations

Performance

EngineSpeedQualityLanguagesNotes
Edge TTSFast⭐⭐⭐⭐⭐100+Best quality, requires internet
SileroVery Fast⭐⭐⭐Russian onlyOffline, robotic
XTTSSlow⭐⭐⭐⭐⭐16Voice cloning, GPU recommended

Troubleshooting

"No module named 'soundfile'"

pip install soundfile

"TorchCodec is required"

This is already patched in the code. If you still see it, update PyTorch:

pip install --upgrade torch torchaudio

Silero model download fails

The script will auto-download on first run (~40MB). Check your internet connection.

XTTS out of memory

Use CPU mode or reduce video length. For long videos, split into segments.

Poor voice quality with Silero

Use Edge TTS or XTTS instead. Silero is designed for speed, not quality.

Ollama Integration Features

  1. Privacy: Your transcripts and translations never leave your local machine.
  2. Custom Context: LLMs can handle nuances, slang, and technical terms better than basic translators.
  3. Cost: 100% free with no character limits or subscription fees.
  4. Offline Workflow: Combined with Silero or XTTS, you can dub videos without an active internet connection.
FeatureGoogle TranslateOllama (Local LLM)
SpeedInstantDepends on your GPU/RAM
SetupZero setupRequires model download
InternetRequiredNot required
QualityLiteral / StandardContextual / Natural

Technical Details

Processing Pipeline

  1. Extract Audio: FFmpeg extracts mono 16kHz WAV
  2. Transcribe: Whisper "base" model transcribes with timestamps
  3. Translate: Google Translate API translates segments
  4. Synthesize: TTS engine generates speech for each subtitle
  5. Merge: FFmpeg mixes original (20%) + dubbed (150%) audio with video

Audio Mixing

  • Original audio: 20% volume (background)
  • Dubbed audio: 150% volume (foreground)
  • Output: AAC 128kbps, video copied without re-encoding

License

MIT License - see LICENSE file

Credits

Contributing

Issues and pull requests welcome!

About

Automatic video translator and dubber using Whisper, XTTS v2 for voice cloning, and Ollama for local LLM translation. Supports 100+ languages.

Topics

Resources

Stars

24 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - shyhirt/AutoDub: Automatic video translator and dubber using Whisper, XTTS v2 for voice cloning, and Ollama for local LLM translation. Supports 100+ languages. · GitHub
Skip to content

Latest commit

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

AutoDub - Automatic Video Translation & Dubbing

Automatically transcribe, translate, and dub videos into different languages using AI-powered text-to-speech.

Features

  • 🎙️ Speech Recognition: Transcribe audio using OpenAI Whisper
  • 🌍 Translation: Translate to 100+ languages via Google Translate
  • 🗣️ Three TTS Engines:
    • Edge TTS: High-quality Microsoft voices (recommended)
    • Silero: Fast Russian TTS (offline after first download)
    • XTTS: Voice cloning from 6-10 second samples
  • 🎬 Video Preservation: Keeps original video, mixes original audio (20%) with dubbed audio (150%)
  • 📝 Subtitle Generation: Creates SRT files for translated text

Requirements

System Dependencies

# Fedora/RHEL
sudo dnf install ffmpeg python3.10 python3.10-devel
# Ubuntu/Debian
sudo apt install ffmpeg python3.10 python3.10-devel
# macOS
brew install ffmpeg

Python Dependencies

python3 -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
pip install torch torchaudio --index-url https://download.pytorch.org/whl/cpu
pip install openai-whisper pysrt edge-tts deep-translator soundfile tqdm
pip install TTS # Only needed for XTTS voice cloning

Local Translation with Ollama (Optional)

AutoDub now supports fully offline translation using Ollama. This is ideal for privacy, avoiding API limits, and achieving more context-aware translations.

1. Install Ollama

For Linux (Fedora/Ubuntu/etc.):

curl -fsSL [https://ollama.com/install.sh](https://ollama.com/install.sh) | sh
ollama pull llama3

Quick Start

Here is the concise guide on how to get started using your setup.sh script, formatted in Markdown: 🚀 Quick Start Guide

Follow these three steps to set up and start dubbing your videos:

  1. Prepare Files

Ensure you have the following files in your project directory:

setup.sh (The installer)
autodub_v4_1.py (The main engine)
install.txt (List of dependencies)
  1. Run Installation

Open your terminal in the project folder and execute:

chmod +x setup.sh && ./setup.sh

Basic Usage (Edge TTS - Recommended)

# Dub to Russian (default)
python autodub.py video.mp4
# Dub to English
python autodub.py video.mp4 --target_lang en
# Dub to German
python autodub.py video.mp4 --target_lang de

Silero TTS (Faster, Russian only)

# Default voice (aidar)
python autodub.py video.mp4 --tts silero
# Female voice
python autodub.py video.mp4 --tts silero --silero_voice xenia
# Available voices: aidar, baya, kseniya, xenia, eugene

XTTS Voice Cloning (Most Natural)

# Requires 6-10 second clean voice sample
python autodub.py video.mp4 --tts xtts --ref_voice my_voice.wav --target_lang en

Ollama translator

# Use Ollama with default llama3 model
./run.sh video.mp4 --translator ollama
# Use a specific model (e.g., Mistral)
./run.sh video.mp4 --translator ollama --ollama_model mistral 

Command-Line Options

positional arguments:
video Input video file
options:
-h, --help Show help message
--tts {edge,silero,xtts}
TTS engine (default: edge)
--target_lang LANG Target language code (default: ru)
Supports: ru, en, de, fr, es, it, pt, ja, zh, etc.
--silero_voice {aidar,baya,kseniya,xenia,eugene}
Silero voice for Russian (default: aidar)
--ref_voice FILE Reference WAV for XTTS voice cloning
--keep-temp Keep temporary files after processing

Supported Languages

Edge TTS supports 100+ languages. Common codes:

  • ru - Russian
  • en - English
  • de - German
  • fr - French
  • es - Spanish
  • it - Italian
  • pt - Portuguese
  • ja - Japanese
  • zh - Chinese

Full list: https://speech.microsoft.com/portal/voicegallery

Output

The script generates:

  • {video}_dubbed.mp4 - Video with dubbed audio
  • {video}_{lang}.srt - Subtitle file with translations

Performance

EngineSpeedQualityLanguagesNotes
Edge TTSFast⭐⭐⭐⭐⭐100+Best quality, requires internet
SileroVery Fast⭐⭐⭐Russian onlyOffline, robotic
XTTSSlow⭐⭐⭐⭐⭐16Voice cloning, GPU recommended

Troubleshooting

"No module named 'soundfile'"

pip install soundfile

"TorchCodec is required"

This is already patched in the code. If you still see it, update PyTorch:

pip install --upgrade torch torchaudio

Silero model download fails

The script will auto-download on first run (~40MB). Check your internet connection.

XTTS out of memory

Use CPU mode or reduce video length. For long videos, split into segments.

Poor voice quality with Silero

Use Edge TTS or XTTS instead. Silero is designed for speed, not quality.

Ollama Integration Features

  1. Privacy: Your transcripts and translations never leave your local machine.
  2. Custom Context: LLMs can handle nuances, slang, and technical terms better than basic translators.
  3. Cost: 100% free with no character limits or subscription fees.
  4. Offline Workflow: Combined with Silero or XTTS, you can dub videos without an active internet connection.
FeatureGoogle TranslateOllama (Local LLM)
SpeedInstantDepends on your GPU/RAM
SetupZero setupRequires model download
InternetRequiredNot required
QualityLiteral / StandardContextual / Natural

Technical Details

Processing Pipeline

  1. Extract Audio: FFmpeg extracts mono 16kHz WAV
  2. Transcribe: Whisper "base" model transcribes with timestamps
  3. Translate: Google Translate API translates segments
  4. Synthesize: TTS engine generates speech for each subtitle
  5. Merge: FFmpeg mixes original (20%) + dubbed (150%) audio with video

Audio Mixing

  • Original audio: 20% volume (background)
  • Dubbed audio: 150% volume (foreground)
  • Output: AAC 128kbps, video copied without re-encoding

License

MIT License - see LICENSE file

Credits

Contributing

Issues and pull requests welcome!

About

Automatic video translator and dubber using Whisper, XTTS v2 for voice cloning, and Ollama for local LLM translation. Supports 100+ languages.

Topics

Resources

Stars

24 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - shyhirt/AutoDub: Automatic video translator and dubber using Whisper, XTTS v2 for voice cloning, and Ollama for local LLM translation. Supports 100+ languages. · GitHub
Skip to content

Latest commit

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

AutoDub - Automatic Video Translation & Dubbing

Automatically transcribe, translate, and dub videos into different languages using AI-powered text-to-speech.

Features

  • 🎙️ Speech Recognition: Transcribe audio using OpenAI Whisper
  • 🌍 Translation: Translate to 100+ languages via Google Translate
  • 🗣️ Three TTS Engines:
    • Edge TTS: High-quality Microsoft voices (recommended)
    • Silero: Fast Russian TTS (offline after first download)
    • XTTS: Voice cloning from 6-10 second samples
  • 🎬 Video Preservation: Keeps original video, mixes original audio (20%) with dubbed audio (150%)
  • 📝 Subtitle Generation: Creates SRT files for translated text

Requirements

System Dependencies

# Fedora/RHEL
sudo dnf install ffmpeg python3.10 python3.10-devel
# Ubuntu/Debian
sudo apt install ffmpeg python3.10 python3.10-devel
# macOS
brew install ffmpeg

Python Dependencies

python3 -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
pip install torch torchaudio --index-url https://download.pytorch.org/whl/cpu
pip install openai-whisper pysrt edge-tts deep-translator soundfile tqdm
pip install TTS # Only needed for XTTS voice cloning

Local Translation with Ollama (Optional)

AutoDub now supports fully offline translation using Ollama. This is ideal for privacy, avoiding API limits, and achieving more context-aware translations.

1. Install Ollama

For Linux (Fedora/Ubuntu/etc.):

curl -fsSL [https://ollama.com/install.sh](https://ollama.com/install.sh) | sh
ollama pull llama3

Quick Start

Here is the concise guide on how to get started using your setup.sh script, formatted in Markdown: 🚀 Quick Start Guide

Follow these three steps to set up and start dubbing your videos:

  1. Prepare Files

Ensure you have the following files in your project directory:

setup.sh (The installer)
autodub_v4_1.py (The main engine)
install.txt (List of dependencies)
  1. Run Installation

Open your terminal in the project folder and execute:

chmod +x setup.sh && ./setup.sh

Basic Usage (Edge TTS - Recommended)

# Dub to Russian (default)
python autodub.py video.mp4
# Dub to English
python autodub.py video.mp4 --target_lang en
# Dub to German
python autodub.py video.mp4 --target_lang de

Silero TTS (Faster, Russian only)

# Default voice (aidar)
python autodub.py video.mp4 --tts silero
# Female voice
python autodub.py video.mp4 --tts silero --silero_voice xenia
# Available voices: aidar, baya, kseniya, xenia, eugene

XTTS Voice Cloning (Most Natural)

# Requires 6-10 second clean voice sample
python autodub.py video.mp4 --tts xtts --ref_voice my_voice.wav --target_lang en

Ollama translator

# Use Ollama with default llama3 model
./run.sh video.mp4 --translator ollama
# Use a specific model (e.g., Mistral)
./run.sh video.mp4 --translator ollama --ollama_model mistral 

Command-Line Options

positional arguments:
video Input video file
options:
-h, --help Show help message
--tts {edge,silero,xtts}
TTS engine (default: edge)
--target_lang LANG Target language code (default: ru)
Supports: ru, en, de, fr, es, it, pt, ja, zh, etc.
--silero_voice {aidar,baya,kseniya,xenia,eugene}
Silero voice for Russian (default: aidar)
--ref_voice FILE Reference WAV for XTTS voice cloning
--keep-temp Keep temporary files after processing

Supported Languages

Edge TTS supports 100+ languages. Common codes:

  • ru - Russian
  • en - English
  • de - German
  • fr - French
  • es - Spanish
  • it - Italian
  • pt - Portuguese
  • ja - Japanese
  • zh - Chinese

Full list: https://speech.microsoft.com/portal/voicegallery

Output

The script generates:

  • {video}_dubbed.mp4 - Video with dubbed audio
  • {video}_{lang}.srt - Subtitle file with translations

Performance

EngineSpeedQualityLanguagesNotes
Edge TTSFast⭐⭐⭐⭐⭐100+Best quality, requires internet
SileroVery Fast⭐⭐⭐Russian onlyOffline, robotic
XTTSSlow⭐⭐⭐⭐⭐16Voice cloning, GPU recommended

Troubleshooting

"No module named 'soundfile'"

pip install soundfile

"TorchCodec is required"

This is already patched in the code. If you still see it, update PyTorch:

pip install --upgrade torch torchaudio

Silero model download fails

The script will auto-download on first run (~40MB). Check your internet connection.

XTTS out of memory

Use CPU mode or reduce video length. For long videos, split into segments.

Poor voice quality with Silero

Use Edge TTS or XTTS instead. Silero is designed for speed, not quality.

Ollama Integration Features

  1. Privacy: Your transcripts and translations never leave your local machine.
  2. Custom Context: LLMs can handle nuances, slang, and technical terms better than basic translators.
  3. Cost: 100% free with no character limits or subscription fees.
  4. Offline Workflow: Combined with Silero or XTTS, you can dub videos without an active internet connection.
FeatureGoogle TranslateOllama (Local LLM)
SpeedInstantDepends on your GPU/RAM
SetupZero setupRequires model download
InternetRequiredNot required
QualityLiteral / StandardContextual / Natural

Technical Details

Processing Pipeline

  1. Extract Audio: FFmpeg extracts mono 16kHz WAV
  2. Transcribe: Whisper "base" model transcribes with timestamps
  3. Translate: Google Translate API translates segments
  4. Synthesize: TTS engine generates speech for each subtitle
  5. Merge: FFmpeg mixes original (20%) + dubbed (150%) audio with video

Audio Mixing

  • Original audio: 20% volume (background)
  • Dubbed audio: 150% volume (foreground)
  • Output: AAC 128kbps, video copied without re-encoding

License

MIT License - see LICENSE file

Credits

Contributing

Issues and pull requests welcome!

About

Automatic video translator and dubber using Whisper, XTTS v2 for voice cloning, and Ollama for local LLM translation. Supports 100+ languages.

Topics

Resources

Stars

24 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - shyhirt/AutoDub: Automatic video translator and dubber using Whisper, XTTS v2 for voice cloning, and Ollama for local LLM translation. Supports 100+ languages. · GitHub
Skip to content

Latest commit

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

AutoDub - Automatic Video Translation & Dubbing

Automatically transcribe, translate, and dub videos into different languages using AI-powered text-to-speech.

Features

  • 🎙️ Speech Recognition: Transcribe audio using OpenAI Whisper
  • 🌍 Translation: Translate to 100+ languages via Google Translate
  • 🗣️ Three TTS Engines:
    • Edge TTS: High-quality Microsoft voices (recommended)
    • Silero: Fast Russian TTS (offline after first download)
    • XTTS: Voice cloning from 6-10 second samples
  • 🎬 Video Preservation: Keeps original video, mixes original audio (20%) with dubbed audio (150%)
  • 📝 Subtitle Generation: Creates SRT files for translated text

Requirements

System Dependencies

# Fedora/RHEL
sudo dnf install ffmpeg python3.10 python3.10-devel
# Ubuntu/Debian
sudo apt install ffmpeg python3.10 python3.10-devel
# macOS
brew install ffmpeg

Python Dependencies

python3 -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
pip install torch torchaudio --index-url https://download.pytorch.org/whl/cpu
pip install openai-whisper pysrt edge-tts deep-translator soundfile tqdm
pip install TTS # Only needed for XTTS voice cloning

Local Translation with Ollama (Optional)

AutoDub now supports fully offline translation using Ollama. This is ideal for privacy, avoiding API limits, and achieving more context-aware translations.

1. Install Ollama

For Linux (Fedora/Ubuntu/etc.):

curl -fsSL [https://ollama.com/install.sh](https://ollama.com/install.sh) | sh
ollama pull llama3

Quick Start

Here is the concise guide on how to get started using your setup.sh script, formatted in Markdown: 🚀 Quick Start Guide

Follow these three steps to set up and start dubbing your videos:

  1. Prepare Files

Ensure you have the following files in your project directory:

setup.sh (The installer)
autodub_v4_1.py (The main engine)
install.txt (List of dependencies)
  1. Run Installation

Open your terminal in the project folder and execute:

chmod +x setup.sh && ./setup.sh

Basic Usage (Edge TTS - Recommended)

# Dub to Russian (default)
python autodub.py video.mp4
# Dub to English
python autodub.py video.mp4 --target_lang en
# Dub to German
python autodub.py video.mp4 --target_lang de

Silero TTS (Faster, Russian only)

# Default voice (aidar)
python autodub.py video.mp4 --tts silero
# Female voice
python autodub.py video.mp4 --tts silero --silero_voice xenia
# Available voices: aidar, baya, kseniya, xenia, eugene

XTTS Voice Cloning (Most Natural)

# Requires 6-10 second clean voice sample
python autodub.py video.mp4 --tts xtts --ref_voice my_voice.wav --target_lang en

Ollama translator

# Use Ollama with default llama3 model
./run.sh video.mp4 --translator ollama
# Use a specific model (e.g., Mistral)
./run.sh video.mp4 --translator ollama --ollama_model mistral 

Command-Line Options

positional arguments:
video Input video file
options:
-h, --help Show help message
--tts {edge,silero,xtts}
TTS engine (default: edge)
--target_lang LANG Target language code (default: ru)
Supports: ru, en, de, fr, es, it, pt, ja, zh, etc.
--silero_voice {aidar,baya,kseniya,xenia,eugene}
Silero voice for Russian (default: aidar)
--ref_voice FILE Reference WAV for XTTS voice cloning
--keep-temp Keep temporary files after processing

Supported Languages

Edge TTS supports 100+ languages. Common codes:

  • ru - Russian
  • en - English
  • de - German
  • fr - French
  • es - Spanish
  • it - Italian
  • pt - Portuguese
  • ja - Japanese
  • zh - Chinese

Full list: https://speech.microsoft.com/portal/voicegallery

Output

The script generates:

  • {video}_dubbed.mp4 - Video with dubbed audio
  • {video}_{lang}.srt - Subtitle file with translations

Performance

EngineSpeedQualityLanguagesNotes
Edge TTSFast⭐⭐⭐⭐⭐100+Best quality, requires internet
SileroVery Fast⭐⭐⭐Russian onlyOffline, robotic
XTTSSlow⭐⭐⭐⭐⭐16Voice cloning, GPU recommended

Troubleshooting

"No module named 'soundfile'"

pip install soundfile

"TorchCodec is required"

This is already patched in the code. If you still see it, update PyTorch:

pip install --upgrade torch torchaudio

Silero model download fails

The script will auto-download on first run (~40MB). Check your internet connection.

XTTS out of memory

Use CPU mode or reduce video length. For long videos, split into segments.

Poor voice quality with Silero

Use Edge TTS or XTTS instead. Silero is designed for speed, not quality.

Ollama Integration Features

  1. Privacy: Your transcripts and translations never leave your local machine.
  2. Custom Context: LLMs can handle nuances, slang, and technical terms better than basic translators.
  3. Cost: 100% free with no character limits or subscription fees.
  4. Offline Workflow: Combined with Silero or XTTS, you can dub videos without an active internet connection.
FeatureGoogle TranslateOllama (Local LLM)
SpeedInstantDepends on your GPU/RAM
SetupZero setupRequires model download
InternetRequiredNot required
QualityLiteral / StandardContextual / Natural

Technical Details

Processing Pipeline

  1. Extract Audio: FFmpeg extracts mono 16kHz WAV
  2. Transcribe: Whisper "base" model transcribes with timestamps
  3. Translate: Google Translate API translates segments
  4. Synthesize: TTS engine generates speech for each subtitle
  5. Merge: FFmpeg mixes original (20%) + dubbed (150%) audio with video

Audio Mixing

  • Original audio: 20% volume (background)
  • Dubbed audio: 150% volume (foreground)
  • Output: AAC 128kbps, video copied without re-encoding

License

MIT License - see LICENSE file

Credits

Contributing

Issues and pull requests welcome!

About

Automatic video translator and dubber using Whisper, XTTS v2 for voice cloning, and Ollama for local LLM translation. Supports 100+ languages.

Topics

Resources

Stars

24 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - shyhirt/AutoDub: Automatic video translator and dubber using Whisper, XTTS v2 for voice cloning, and Ollama for local LLM translation. Supports 100+ languages. · GitHub
Skip to content

Latest commit

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

AutoDub - Automatic Video Translation & Dubbing

Automatically transcribe, translate, and dub videos into different languages using AI-powered text-to-speech.

Features

  • 🎙️ Speech Recognition: Transcribe audio using OpenAI Whisper
  • 🌍 Translation: Translate to 100+ languages via Google Translate
  • 🗣️ Three TTS Engines:
    • Edge TTS: High-quality Microsoft voices (recommended)
    • Silero: Fast Russian TTS (offline after first download)
    • XTTS: Voice cloning from 6-10 second samples
  • 🎬 Video Preservation: Keeps original video, mixes original audio (20%) with dubbed audio (150%)
  • 📝 Subtitle Generation: Creates SRT files for translated text

Requirements

System Dependencies

# Fedora/RHEL
sudo dnf install ffmpeg python3.10 python3.10-devel
# Ubuntu/Debian
sudo apt install ffmpeg python3.10 python3.10-devel
# macOS
brew install ffmpeg

Python Dependencies

python3 -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
pip install torch torchaudio --index-url https://download.pytorch.org/whl/cpu
pip install openai-whisper pysrt edge-tts deep-translator soundfile tqdm
pip install TTS # Only needed for XTTS voice cloning

Local Translation with Ollama (Optional)

AutoDub now supports fully offline translation using Ollama. This is ideal for privacy, avoiding API limits, and achieving more context-aware translations.

1. Install Ollama

For Linux (Fedora/Ubuntu/etc.):

curl -fsSL [https://ollama.com/install.sh](https://ollama.com/install.sh) | sh
ollama pull llama3

Quick Start

Here is the concise guide on how to get started using your setup.sh script, formatted in Markdown: 🚀 Quick Start Guide

Follow these three steps to set up and start dubbing your videos:

  1. Prepare Files

Ensure you have the following files in your project directory:

setup.sh (The installer)
autodub_v4_1.py (The main engine)
install.txt (List of dependencies)
  1. Run Installation

Open your terminal in the project folder and execute:

chmod +x setup.sh && ./setup.sh

Basic Usage (Edge TTS - Recommended)

# Dub to Russian (default)
python autodub.py video.mp4
# Dub to English
python autodub.py video.mp4 --target_lang en
# Dub to German
python autodub.py video.mp4 --target_lang de

Silero TTS (Faster, Russian only)

# Default voice (aidar)
python autodub.py video.mp4 --tts silero
# Female voice
python autodub.py video.mp4 --tts silero --silero_voice xenia
# Available voices: aidar, baya, kseniya, xenia, eugene

XTTS Voice Cloning (Most Natural)

# Requires 6-10 second clean voice sample
python autodub.py video.mp4 --tts xtts --ref_voice my_voice.wav --target_lang en

Ollama translator

# Use Ollama with default llama3 model
./run.sh video.mp4 --translator ollama
# Use a specific model (e.g., Mistral)
./run.sh video.mp4 --translator ollama --ollama_model mistral 

Command-Line Options

positional arguments:
video Input video file
options:
-h, --help Show help message
--tts {edge,silero,xtts}
TTS engine (default: edge)
--target_lang LANG Target language code (default: ru)
Supports: ru, en, de, fr, es, it, pt, ja, zh, etc.
--silero_voice {aidar,baya,kseniya,xenia,eugene}
Silero voice for Russian (default: aidar)
--ref_voice FILE Reference WAV for XTTS voice cloning
--keep-temp Keep temporary files after processing

Supported Languages

Edge TTS supports 100+ languages. Common codes:

  • ru - Russian
  • en - English
  • de - German
  • fr - French
  • es - Spanish
  • it - Italian
  • pt - Portuguese
  • ja - Japanese
  • zh - Chinese

Full list: https://speech.microsoft.com/portal/voicegallery

Output

The script generates:

  • {video}_dubbed.mp4 - Video with dubbed audio
  • {video}_{lang}.srt - Subtitle file with translations

Performance

EngineSpeedQualityLanguagesNotes
Edge TTSFast⭐⭐⭐⭐⭐100+Best quality, requires internet
SileroVery Fast⭐⭐⭐Russian onlyOffline, robotic
XTTSSlow⭐⭐⭐⭐⭐16Voice cloning, GPU recommended

Troubleshooting

"No module named 'soundfile'"

pip install soundfile

"TorchCodec is required"

This is already patched in the code. If you still see it, update PyTorch:

pip install --upgrade torch torchaudio

Silero model download fails

The script will auto-download on first run (~40MB). Check your internet connection.

XTTS out of memory

Use CPU mode or reduce video length. For long videos, split into segments.

Poor voice quality with Silero

Use Edge TTS or XTTS instead. Silero is designed for speed, not quality.

Ollama Integration Features

  1. Privacy: Your transcripts and translations never leave your local machine.
  2. Custom Context: LLMs can handle nuances, slang, and technical terms better than basic translators.
  3. Cost: 100% free with no character limits or subscription fees.
  4. Offline Workflow: Combined with Silero or XTTS, you can dub videos without an active internet connection.
FeatureGoogle TranslateOllama (Local LLM)
SpeedInstantDepends on your GPU/RAM
SetupZero setupRequires model download
InternetRequiredNot required
QualityLiteral / StandardContextual / Natural

Technical Details

Processing Pipeline

  1. Extract Audio: FFmpeg extracts mono 16kHz WAV
  2. Transcribe: Whisper "base" model transcribes with timestamps
  3. Translate: Google Translate API translates segments
  4. Synthesize: TTS engine generates speech for each subtitle
  5. Merge: FFmpeg mixes original (20%) + dubbed (150%) audio with video

Audio Mixing

  • Original audio: 20% volume (background)
  • Dubbed audio: 150% volume (foreground)
  • Output: AAC 128kbps, video copied without re-encoding

License

MIT License - see LICENSE file

Credits

Contributing

Issues and pull requests welcome!

About

Automatic video translator and dubber using Whisper, XTTS v2 for voice cloning, and Ollama for local LLM translation. Supports 100+ languages.

Topics

Resources

Stars

24 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages