TTS Channel: Qwen-TTS
Supported languages: Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian
Qwen-TTS is an advanced speech synthesis technology from Alibaba's Tongyi Qianwen team, capable of converting text into highly realistic, natural-sounding human voice. A key highlight is its ability to automatically adjust speech rhythm and emotion based on text content.
pyVideoTrans supports two modes of Qwen3-TTS:
- Built-in local version (offline): Bundled with the software, no internet required, uses a fixed 1.7B model
- Alibaba Bailian API (online): Called via Alibaba Cloud API, requires internet and an API Key
Qwen3-TTS Local Built-in (Offline)
Prerequisites
| Requirement | Details |
|---|---|
| pyVideoTrans version | ≥ v3.97 |
| Model size | ~8GB (auto-downloaded on first use) |
| Hardware | NVIDIA GPU recommended |
Step 1: Confirm Version
Make sure you've upgraded to v3.97 or higher. The built-in version uses a fixed 1.7B model.
Step 2: Download Model (auto on first use)
The first launch will automatically download both the Base and CustomVoice models, totaling about 8GB. Please be patient.
Manual Download (optional)
If automatic download is too slow:
Open the
modelsfolder in the software directory and create two new folders:models--Qwen--Qwen3-TTS-12Hz-1.7B-Basemodels--Qwen--Qwen3-TTS-12Hz-1.7B-CustomVoice
Open the Base model download page and download all files into
models/models--Qwen--Qwen3-TTS-12Hz-1.7B-BaseOpen the CustomVoice model download page and download all files into
models/models--Qwen--Qwen3-TTS-12Hz-1.7B-CustomVoice
As shown below:



Step 3: Configure Reference Audio
Used for cloning a voice based on a 3–10 second reference audio clip.
Path: Menu → Tools → TTS Settings → Qwen-tts (Local)
Enter the reference audio and its corresponding text content, one entry per line.
Format
audio_filename.wav#text corresponding to the audioExample
n10.wav#You say all is empty, yet you keep your eyes closed. Open them and look at me — I don't believe your eyes are truly empty.Place the n10.wav file in the f5-tts folder under the software directory, then enter the speech text after the # symbol.


Step 4: Voice Style Prompt (optional)
When using the built-in Vivian, Uncle_fu, or Sohee preset voices, you can enter a prompt to control the speech style.
Path: Menu → Tools → TTS Settings → Qwen-tts (Local)
Enter a short prompt in the "Prompt" text box, for example:
Use an angry, furious toneThis prompt will be automatically applied when using the built-in voices.
Qwen3-TTS Alibaba Bailian API (Online)
The qwen3-tts model supports 10 languages and multiple Chinese dialects Model name:
qwen3-tts-flashClick here for available voices and supported language details
Step 1: Obtain and Configure API Key
- Click this link to access the Alibaba Cloud Bailian console: https://bailian.console.aliyun.com/?tab=model#/api-key

- Log in to your Alibaba Cloud account (register one if you don't have an account)
- On the API-KEY management page, click "Create API-KEY". The system will generate a string starting with
sk-. Copy this string - In pyVideoTrans, go to the top menu bar: TTS Settings → Qwen TTS

- In the "Qwen3 TTS" configuration window, paste the API KEY into the input box. Click "Test" to preview the effect — if you can hear audio, the configuration is successful. Finally, click "Save".

Step 2: Using Qwen3-TTS in Video Translation
After configuration, select "Qwen3 TTS" from the "TTS Channel" dropdown on the main screen, then pick your preferred voice from the "Dubbing Character" menu.
- Cherry: Standard female voice
- Sunny: Sichuan dialect
- Dylan: Beijing dialect
- See the full list below

Step 3: Using in Batch and Multi-Character Dubbing
Qwen-TTS's powerful features also work for batch processing:
- Batch subtitle dubbing: In the "Batch Dubbing" interface, select "Qwen TTS" and your desired character from the "TTS Channel" dropdown
- Multi-character subtitle dubbing: In the "Multi-Character Dubbing" feature, assign different Qwen-TTS voices to different characters

Available Voices
Here is the full list of voices supported by Qwen3-TTS (online version):
| Chinese Name | English Code | Type |
|---|---|---|
| Qianyue (Cherry) | Cherry | Standard female |
| Suyao (Serena) | Serena | Standard female |
| Chenxu (Ethan) | Ethan | Standard male |
| Qianxue (Chelsie) | Chelsie | Standard female |
| Motu (Momo) | Momo | Standard female |
| Shisan (Vivian) | Vivian | Standard female |
| Yuebai (Moon) | Moon | Standard female |
| Siyue (Maia) | Maia | Standard female |
| Kai | Kai | Standard male |
| Buciyu (Nofish) | Nofish | Standard male |
| Mengbao (Bella) | Bella | Child voice |
| Jennifer | Jennifer | English female |
| Tiancha (Ryan) | Ryan | English male |
| Katerina | Katerina | Russian female |
| Aiden | Aiden | English male |
| Cangmingzi (Eldric Sage) | Eldric Sage | English male |
| Guaixiaomei (Mia) | Mia | Standard female |
| Shaxiaomi (Mochi) | Mochi | Standard female |
| Yanzhengying (Bellona) | Bellona | Standard female |
| Tianshu (Vincent) | Vincent | Standard male |
| Mengxiaoji (Bunny) | Bunny | Standard female |
| Awen (Neil) | Neil | Standard male |
| Mojiangshi (Elias) | Elias | Standard male |
| Xudaye (Arthur) | Arthur | Standard male |
| Linjiameimei (Nini) | Nini | Standard female |
| Guipopó (Ebona) | Ebona | Standard female |
| Xiaowan (Seren) | Seren | Standard female |
| Wanpixiaohai (Pip) | Pip | Child voice |
| Shaonv Ayue (Stella) | Stella | Standard female |
| Bodoga (Bodega) | Bodega | Standard male |
| Suonisha (Sonrisa) | Sonrisa | Standard female |
| Aleke (Alek) | Alek | Standard male |
| Duoerqie (Dolce) | Dolce | Standard female |
| Suxi (Sohee) | Sohee | Korean female |
| Xiaoyexing (Ono Anna) | Ono Anna | Japanese female |
| Laien (Lenn) | Lenn | Standard male |
| Aimi'eran (Emilien) | Emilien | French male |
| Andre | Andre | Standard male |
| Ladi'ao Ge'er (Radio Gol) | Radio Gol | Standard male |
| Shanghai - Azhen (Jada) | Jada | Shanghai dialect female |
| Beijing - Xiaodong (Dylan) | Dylan | Beijing dialect male |
| Nanjing - Laoli (Li) | Li | Nanjing dialect male |
| Shaanxi - Qinchuan (Marcus) | Marcus | Shaanxi dialect male |
| Minnan - Ajie (Roy) | Roy | Minnan dialect male |
| Tianjin - Li Peter (Peter) | Peter | Tianjin dialect male |
| Sichuan - Qing'er (Sunny) | Sunny | Sichuan dialect female |
| Sichuan - Chengchuan (Eric) | Eric | Sichuan dialect male |
| Cantonese - Aqiang (Rocky) | Rocky | Cantonese male |
| Cantonese - Aqing (Kiki) | Kiki | Cantonese female |
Troubleshooting
1. Local Version Model Download Is Slow
The first launch downloads about 8GB of model files. Be patient. For faster download, use the manual method described above to download from HuggingFace.
2. API Version Shows AuthenticationError
The API KEY is invalid or expired. Log back into the Alibaba Cloud Bailian console to obtain a new API KEY.
3. Dubbing Sounds Unnatural
- Try a different voice
- For the local version, try entering a voice style prompt
- Ensure the reference audio is high quality (clear pronunciation, no noise)
