Dubbing Channels
Text-to-Speech (TTS) is the third step of video translation — converting translated subtitle text into spoken audio. pyVideoTrans supports 30+ dubbing channels.
Ready-to-Use (Free)
No complex configuration needed — ideal for beginners.
| Provider | Description | Rating |
|---|---|---|
| Edge-TTS (Free) | Microsoft's free interface; natural voice; supports all languages | ⭐⭐⭐ Default recommendation |
| gTTS (Free) | Google TTS; basic quality | ⭐⭐ |
⚠️ Edge-TTS may trigger rate limiting with heavy short-term usage. It is recommended to set concurrency to 1 and pause to 5–10 seconds in More Settings.
Local Built-in (Free)
Models are downloaded automatically on first use and run completely offline.
| Provider | Description | GPU Accelerated | Rating |
|---|---|---|---|
| Qwen3-TTS (Local Built-in) | Alibaba open-source; supports Chinese, English, Japanese, Korean, etc. | ✅ | ⭐⭐⭐ Recommended |
| MOSS-TTS-Nano (Local Built-in) | Supports 20 languages | ❌ | ⭐⭐ |
| Piper (Local Built-in) | Lightweight; supports 20 languages | ❌ | ⭐⭐ |
| VITS (Local Built-in) | Chinese-English dubbing | ❌ | ⭐⭐ |
| Supertonic3 (Local Built-in) | English, Korean, Spanish, French dubbing | ❌ | ⭐⭐ |
| ChatterBox (Local Built-in) | 22 languages; high quality | ✅ | ⭐⭐⭐ Recommended |
Professional Cloud Services (API Key Required)
| Provider | Description | Rating |
|---|---|---|
| Azure TTS | Microsoft's professional speech service | ⭐⭐⭐ |
| OpenAI TTS | Leading voice technology | ⭐⭐⭐ |
| ByteDance TTS 2.0 | Natural Chinese pronunciation | ⭐⭐⭐ |
| Alibaba Qwen-TTS | Alibaba Cloud speech synthesis | ⭐⭐⭐ |
| Gemini TTS | Google TTS | ⭐⭐ |
| Elevenlabs.io | AI audio technology company | ⭐⭐⭐ |
| 302.AI | Aggregation platform | ⭐⭐ |
| Minimaxi | Requires top-up to use | ⭐⭐ |
| Xiaomi TTS | Xiaomi AI open platform | ⭐⭐ |
| X.AI TTS | x.ai platform | ⭐⭐ |
Local Deployment (Advanced)
| Provider | Description | Voice Cloning | Rating |
|---|---|---|---|
| OmniVoice-TTS | Supports nearly all languages | ✅ | ⭐⭐⭐ Recommended |
| GPT-SoVITS | Clone with minimal audio samples | ✅ | ⭐⭐⭐ Recommended |
| F5-TTS | Chinese-English cloning | ✅ | ⭐⭐⭐ Recommended |
| Index-TTS | Chinese-English cloning | ✅ | ⭐⭐⭐ Recommended |
| Confucius-TTS | 14 languages | ✅ | ⭐⭐⭐ |
| VoxCPM-TTS | 10+ languages | ✅ | ⭐⭐⭐ |
| CosyVoice | Chinese, English, Japanese, Korean, and 10+ others | ✅ | ⭐⭐ |
| ChatTTS | Chinese and English | — | ⭐⭐ |
| Fish-TTS | All built-in languages | — | ⭐ |
| Kokoro-TTS | Chinese, English, Korean, Italian, Portuguese, German, French, Hindi | — | ⭐ |
| Spark-TTS | English | ✅ | ⭐⭐ |
| Dia-TTS | English | ✅ | ⭐⭐ |
| clone-voice | No longer maintained | ✅ | ⭐ |
Using Reference Audio
Voice cloning providers require a reference audio file. Place WAV files in the f5-tts/ directory with the format filename.wav#spoken text in the audio.
For details, see Voice Cloning & Multi-Role Dubbing
