Skip to content

TTS Channel: Qwen-TTS

Supported languages: Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian

Qwen-TTS is an advanced speech synthesis technology from Alibaba's Tongyi Qianwen team, capable of converting text into highly realistic, natural-sounding human voice. A key highlight is its ability to automatically adjust speech rhythm and emotion based on text content.

pyVideoTrans supports two modes of Qwen3-TTS:

  • Built-in local version (offline): Bundled with the software, no internet required, uses a fixed 1.7B model
  • Alibaba Bailian API (online): Called via Alibaba Cloud API, requires internet and an API Key

Qwen3-TTS Local Built-in (Offline)

Prerequisites

RequirementDetails
pyVideoTrans version≥ v3.97
Model size~8GB (auto-downloaded on first use)
HardwareNVIDIA GPU recommended

Step 1: Confirm Version

Make sure you've upgraded to v3.97 or higher. The built-in version uses a fixed 1.7B model.

Step 2: Download Model (auto on first use)

The first launch will automatically download both the Base and CustomVoice models, totaling about 8GB. Please be patient.

Manual Download (optional)

If automatic download is too slow:

  1. Open the models folder in the software directory and create two new folders:

    • models--Qwen--Qwen3-TTS-12Hz-1.7B-Base
    • models--Qwen--Qwen3-TTS-12Hz-1.7B-CustomVoice
  2. Open the Base model download page and download all files into models/models--Qwen--Qwen3-TTS-12Hz-1.7B-Base

  3. Open the CustomVoice model download page and download all files into models/models--Qwen--Qwen3-TTS-12Hz-1.7B-CustomVoice

As shown below:

Step 3: Configure Reference Audio

Used for cloning a voice based on a 3–10 second reference audio clip.

Path: Menu → Tools → TTS Settings → Qwen-tts (Local)

Enter the reference audio and its corresponding text content, one entry per line.

Format

audio_filename.wav#text corresponding to the audio

Example

n10.wav#You say all is empty, yet you keep your eyes closed. Open them and look at me — I don't believe your eyes are truly empty.

Place the n10.wav file in the f5-tts folder under the software directory, then enter the speech text after the # symbol.

Step 4: Voice Style Prompt (optional)

When using the built-in Vivian, Uncle_fu, or Sohee preset voices, you can enter a prompt to control the speech style.

Path: Menu → Tools → TTS Settings → Qwen-tts (Local)

Enter a short prompt in the "Prompt" text box, for example:

Use an angry, furious tone

This prompt will be automatically applied when using the built-in voices.


Qwen3-TTS Alibaba Bailian API (Online)

The qwen3-tts model supports 10 languages and multiple Chinese dialects Model name: qwen3-tts-flashClick here for available voices and supported language details

Step 1: Obtain and Configure API Key

  1. Click this link to access the Alibaba Cloud Bailian console: https://bailian.console.aliyun.com/?tab=model#/api-key

  1. Log in to your Alibaba Cloud account (register one if you don't have an account)
  2. On the API-KEY management page, click "Create API-KEY". The system will generate a string starting with sk-. Copy this string
  3. In pyVideoTrans, go to the top menu bar: TTS Settings → Qwen TTS

  1. In the "Qwen3 TTS" configuration window, paste the API KEY into the input box. Click "Test" to preview the effect — if you can hear audio, the configuration is successful. Finally, click "Save".

Step 2: Using Qwen3-TTS in Video Translation

After configuration, select "Qwen3 TTS" from the "TTS Channel" dropdown on the main screen, then pick your preferred voice from the "Dubbing Character" menu.

  • Cherry: Standard female voice
  • Sunny: Sichuan dialect
  • Dylan: Beijing dialect
  • See the full list below

Step 3: Using in Batch and Multi-Character Dubbing

Qwen-TTS's powerful features also work for batch processing:

  • Batch subtitle dubbing: In the "Batch Dubbing" interface, select "Qwen TTS" and your desired character from the "TTS Channel" dropdown
  • Multi-character subtitle dubbing: In the "Multi-Character Dubbing" feature, assign different Qwen-TTS voices to different characters


Available Voices

Here is the full list of voices supported by Qwen3-TTS (online version):

Chinese NameEnglish CodeType
Qianyue (Cherry)CherryStandard female
Suyao (Serena)SerenaStandard female
Chenxu (Ethan)EthanStandard male
Qianxue (Chelsie)ChelsieStandard female
Motu (Momo)MomoStandard female
Shisan (Vivian)VivianStandard female
Yuebai (Moon)MoonStandard female
Siyue (Maia)MaiaStandard female
KaiKaiStandard male
Buciyu (Nofish)NofishStandard male
Mengbao (Bella)BellaChild voice
JenniferJenniferEnglish female
Tiancha (Ryan)RyanEnglish male
KaterinaKaterinaRussian female
AidenAidenEnglish male
Cangmingzi (Eldric Sage)Eldric SageEnglish male
Guaixiaomei (Mia)MiaStandard female
Shaxiaomi (Mochi)MochiStandard female
Yanzhengying (Bellona)BellonaStandard female
Tianshu (Vincent)VincentStandard male
Mengxiaoji (Bunny)BunnyStandard female
Awen (Neil)NeilStandard male
Mojiangshi (Elias)EliasStandard male
Xudaye (Arthur)ArthurStandard male
Linjiameimei (Nini)NiniStandard female
Guipopó (Ebona)EbonaStandard female
Xiaowan (Seren)SerenStandard female
Wanpixiaohai (Pip)PipChild voice
Shaonv Ayue (Stella)StellaStandard female
Bodoga (Bodega)BodegaStandard male
Suonisha (Sonrisa)SonrisaStandard female
Aleke (Alek)AlekStandard male
Duoerqie (Dolce)DolceStandard female
Suxi (Sohee)SoheeKorean female
Xiaoyexing (Ono Anna)Ono AnnaJapanese female
Laien (Lenn)LennStandard male
Aimi'eran (Emilien)EmilienFrench male
AndreAndreStandard male
Ladi'ao Ge'er (Radio Gol)Radio GolStandard male
Shanghai - Azhen (Jada)JadaShanghai dialect female
Beijing - Xiaodong (Dylan)DylanBeijing dialect male
Nanjing - Laoli (Li)LiNanjing dialect male
Shaanxi - Qinchuan (Marcus)MarcusShaanxi dialect male
Minnan - Ajie (Roy)RoyMinnan dialect male
Tianjin - Li Peter (Peter)PeterTianjin dialect male
Sichuan - Qing'er (Sunny)SunnySichuan dialect female
Sichuan - Chengchuan (Eric)EricSichuan dialect male
Cantonese - Aqiang (Rocky)RockyCantonese male
Cantonese - Aqing (Kiki)KikiCantonese female

Troubleshooting

1. Local Version Model Download Is Slow

The first launch downloads about 8GB of model files. Be patient. For faster download, use the manual method described above to download from HuggingFace.

2. API Version Shows AuthenticationError

The API KEY is invalid or expired. Log back into the Alibaba Cloud Bailian console to obtain a new API KEY.

3. Dubbing Sounds Unnatural

  • Try a different voice
  • For the local version, try entering a voice style prompt
  • Ensure the reference audio is high quality (clear pronunciation, no noise)