Nodes/comfyui-arkennemasis/arkennemasis Qwen3-TTS (voice clone)
ComfyUI Node

arkennemasis Qwen3-TTS (voice clone)

Qwen3-TTS, running locally on the GPU with no account. Wire a clip into `reference_audio` and it speaks your text in that voice.

By Hishamahmer·Created about a month ago·Updated 9 days ago· 5
arkennemasis Qwen3-TTS (voice clone)
  • reference_audio
  • audio
  • report
text
model
language
seed0
reference_text
timeout_seconds180
Categoryarkennemasis/Audio

Inputs (7)

NameTypeDefaultDescription
textSTRINGWhat to say.
modelCOMBOA folder under ComfyUI/models/qwen-tts. Voice cloning needs a *Base* model.
languageCOMBO'Auto' lets the model decide from the text.
seedINT00–18446744073709550000
reference_audiooptAUDIOThe voice to clone. 5-30 seconds of clean speech is plenty. A *Base* model REQUIRES this.
reference_textoptSTRINGA transcript of the reference clip. Supplying it gives a closer clone; leaving it blank uses x-vector-only mode, which still works.
timeout_secondsoptINT18060–3600Per attempt, and there are 3 attempts with fresh seeds. A 24-word line takes about 30 s including the ~2.5 GB weight load, so 180 s is six times over — it is a runaway detector, not a budget. It was 600 s, which meant one looping line cost 10 minutes before failing.

Outputs (2)

NameTypeDescription
audioAUDIO
reportSTRING