Nodes/APZmedia Qwen TTS Nodes/APZmedia: Voice Design
ComfyUI Node

APZmedia: Voice Design

A ComfyUI node in APZmedia/TTS with 7 inputs and 4 outputs.

By APZmedia·Created 5 months ago·Updated 5 months ago· 1
APZmedia: Voice Design
  • model_design
  • model_base
  • voice_prompt
  • reference_audio
  • reference_text
  • voice_description
voice_descriptionWarm female voice, calm and articulate.
reference_textHello, this is a short reference line for voice cloning.
languageEnglish
x_vector_only_modefalse
seed0
CategoryAPZmedia/TTS

Inputs (7)

NameTypeDefaultDescription
model_designQWEN_TTS_MODELLoad a VoiceDesign model (Qwen3-TTS-12Hz-*-VoiceDesign).
model_baseQWEN_TTS_MODELLoad a Base model (Qwen3-TTS-12Hz-*-Base) to build the speaker embedding.
voice_descriptionSTRINGWarm female voice, calm and articulate.Natural-language voice character and delivery style. Examples: 'Deep authoritative male, slightly amused', 'Young energetic voice, nervous and rushing'.
reference_textSTRINGHello, this is a short reference line for voice cloning.Text spoken in the reference clip used to build the embedding. Keep it under 20 words. Transcript must match the audio exactly.
languageCOMBOEnglish11 options: English, Chinese, Japanese, Korean, German, French, +5
x_vector_only_modeBOOLEANfalseTrue: use only the speaker x-vector (ignores ref_text, faster). False: full in-context learning — better voice fidelity but ref_text must be accurate.
seedINT00–2147483647

Outputs (4)

NameTypeDescription
voice_promptVOICE_PROMPT
reference_audioAUDIO
reference_textSTRING
voice_descriptionSTRING