Nodes/ComfyUI-Lyonir-Studio/🐺 Lyonir Qwen3-TTS Voice Clone
ComfyUI Node

🐺 Lyonir Qwen3-TTS Voice Clone

Clone a Voice From a Clip — and Let It Pick Its Own Best Take

By Lyonir·Created 4 days ago·Updated a day ago· 2
🐺 Lyonir Qwen3-TTS Voice Clone
  • ref_audio
  • audio
◄target_textOlá! Esta é uma clonagem de voz em português brasileiro.►
◄model_choice1.7B►
◄deviceauto►
◄precisionbf16►
◄languagePortuguese (Brazil)►
◄seed0►
◄max_new_tokens2048►
◄top_p1.00►
◄top_k50►
◄temperature0.90►
◄repetition_penalty1.05►
◄attentionauto►
◄output_cleanupClean Voice (recommended)►
◄unload_model_after_generatefalse►
◄custom_model_path►
◄voice_instructionDescreva exatamente como a fala deve ser interpretada. Quando preenchida, o Lyonir usa uma rota Instruction-First para maximizar ritmo, emoção, energia, ênfase, pausas e entonação, preservando a identidade da referência.►
◄brazilian_clone_modeNative PT-BR Hybrid (recommended)►
◄official_clone_modeOfficial Qwen ICL Clone (recommended)►
◄ptbr_checkpoint_step15000►
◄download_ptbr_if_missingtrue►
◄ref_text►
◄clone_profileMaximum Similarity►
◄candidate_count4►
◄reference_processingHQ Clean (recommended)►
◄subtalker_temperature0.65►

Voice cloning is the job that actually moved in the last couple of years - the KB's summary of the TTS field is that local closed the cloning gap, and Qwen3-TTS is one of the reasons. 🐺 Lyonir Qwen3-TTS Voice Clone is the reference-driven node of the pack: hand it a recording, hand it a line, get that voice saying your line. Then it quietly does something most clone nodes don't.

What it's for

Zero-shot cloning. No fine-tune, no training run, no waiting. You need a clean clip of the speaker - a few seconds of speech with no music, no room reverb, no overlapping voices - and the node does the rest. Wire its AUDIO output into the pack's video saver and you've got a character who can talk over the clip you just generated.

One expectation to set: cloning gives you identity, not acting. The upstream clone model is built around matching a speaker, and community reports of Qwen3-TTS in ComfyUI are clear that delivery control is not what the clone path is for. So treat the voice_instruction field on this node as exactly what the node description calls it - an "Instruction-First performance guide" that runs alongside a real reference supplying identity. It genuinely helps with pacing and emphasis. It is not a character slider.

How it works

The pack's clone module is explicit about its choices, and they're the right ones. It uses the ICL prompt path - reference audio plus an exact transcript - whenever it can, because that gives higher fidelity than feeding only the speaker embedding. Reference cleanup is deliberately conservative: mono conversion, DC removal, edge-silence trim, peak normalization. No denoising, no EQ, on purpose - those can alter the very timbre you're trying to clone.

Then there's candidate_count, which is the feature worth the download. Instead of generating one take and hoping, the node renders up to six candidates, embeds each one with Qwen's own speaker encoder, compares them by cosine similarity to the reference embedding (plus a lightweight delivery-similarity score against the guide audio), and returns the closest. Four candidates by default. It costs time; it buys consistency. If you're chasing a specific voice, that's a fair trade.

For Brazilian Portuguese there's a second layer, the same one the other two TTS nodes use: a dedicated Brazilian checkpoint acts as the accent and prosody source, so you never get generic Portuguese-accented output. All four brazilian_clone_mode options use it. That checkpoint gets snapshotted from Hugging Face into ComfyUI/models/qwen-tts/fala_pb_checkpoints/ when download_ptbr_if_missing is on.

Inputs and outputs

Required: ref_audio (AUDIO - your reference clip), target_text (multiline), model_choice (1.7B or 0.6B), device, precision, language (12 options). The model_choice tooltip notes that PT-BR Hybrid uses Base 1.7B internally for quality, while 0.6B applies to the official clone path for other languages - which is to say the PT-BR route ignores your 0.6B choice.

The optional pile is where the real knobs live: voice_instruction (the performance guide), brazilian_clone_mode, official_clone_mode - "Official Qwen ICL Clone (recommended)" or "Official Qwen Speaker-Only"; the tooltip explains ICL uses audio plus transcription for higher fidelity while Speaker-Only uses the vocal identity alone - ref_text, clone_profile (default Maximum Similarity), candidate_count (1–6, default 4), reference_processing (HQ Clean recommended, or Original), subtalker_temperature (0.65), plus the usual seed, temperature, top_p, top_k, repetition_penalty, max_new_tokens, attention, output_cleanup, unload_model_after_generate, custom_model_path, ptbr_checkpoint_step, download_ptbr_if_missing.

Output: one audio, typed AUDIO.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/Lyonir/ComfyUI-Lyonir-Studio.git
git clone https://github.com/flybirdxx/ComfyUI-Qwen-TTS.git
python -m pip install -r ComfyUI-Lyonir-Studio/requirements.txt

Weights go under ComfyUI/models/qwen-tts/. Restart ComfyUI and refresh the browser.

Where people trip

PT-BR clone demands a transcript. In Brazilian Portuguese mode, if ref_text is empty the node refuses outright - the error says PT-BR requires ref_text to be the exact transcript of the reference. That's not a suggestion; the accent anchor needs it. Type what the clip actually says, word for word, including "ums" if they're there.

Garbage in, garbage out. A clip with background music, a hissy mic, or two speakers will clone the noise along with the voice. Sixty seconds of clean, close-miked speech beats five minutes of podcast. Keep reference_processing on HQ Clean unless the reference is already pristine.

Cost sneaks up on you. candidate_count at 4 means four full generations plus embedding comparisons per run. Drop it to 1 while you're iterating on the text, then raise it for the take you keep.

Backend missing. If the node complains the Qwen3-TTS backend wasn't found, flybirdxx/ComfyUI-Qwen-TTS isn't installed alongside this pack in custom_nodes, or it failed to import. The pack borrows its bundled qwen_tts rather than pip-installing the standalone package, deliberately, to avoid a Transformers version fight.

Cloning someone and selling the result. Voice likeness is legally live territory - the pack ships NOTICE and COMMERCIAL_LICENSES.md and says to read them before commercial deployment. Read them.

CategoryLyonir Studio/Qwen3-TTS

Inputs (26)

NameTypeDefaultDescription
ref_audioAUDIO—
target_textSTRINGOlá! Esta é uma clonagem de voz em português brasileiro.—
model_choiceCOMBO1.7BPT-BR Hybrid usa Base 1.7B internamente para máxima qualidade. 0.6B é aplicado ao clone oficial dos outros idiomas.
deviceCOMBOauto5 options: auto, cuda, cpu, mps, xpu
precisionCOMBObf163 options: bf16, fp16, fp32
languageCOMBOPortuguese (Brazil)12 options: Auto, Chinese, English, Japanese, Korean, German, +6
seedoptINT00–18446744073709550000—
max_new_tokensoptINT2048256–8192—
top_poptFLOAT1.000–1—
top_koptINT500–200—
temperatureoptFLOAT0.900.1–2—
repetition_penaltyoptFLOAT1.051–2—
attentionoptCOMBOauto5 options: auto, sage_attention, sdpa, eager, flash_attention_2
output_cleanupoptCOMBOClean Voice (recommended)Post-synthesis cleanup for hiss/background noise. Strong Clean is more aggressive; Off returns the raw model output.
unload_model_after_generateoptBOOLEANfalse—
custom_model_pathoptSTRING—
voice_instructionoptSTRINGDescreva exatamente como a fala deve ser interpretada. Quando preenchida, o Lyonir usa uma rota Instruction-First para maximizar ritmo, emoção, energia, ênfase, pausas e entonação, preservando a identidade da referência.—
brazilian_clone_modeoptCOMBONative PT-BR Hybrid (recommended)Usado apenas quando language = Portuguese (Brazil). Todos os modos PT-BR usam o checkpoint brasileiro; nenhum usa o português genérico como fonte final de sotaque.
official_clone_modeoptCOMBOOfficial Qwen ICL Clone (recommended)Usado nos idiomas originais do Qwen. ICL usa áudio + transcrição para maior fidelidade; Speaker-Only usa apenas a identidade vocal.
ptbr_checkpoint_stepoptCOMBO15000Checkpoint 15000 é o recomendado para PT-BR.
download_ptbr_if_missingoptBOOLEANtrue—
ref_textoptSTRING—
clone_profileoptCOMBOMaximum SimilarityAdvanced setting. PT-BR Custom respeita este campo; os presets PT-BR recomendados usam configurações próprias.
candidate_countoptINT41–6—
reference_processingoptCOMBOHQ Clean (recommended)2 options: HQ Clean (recommended), Original
subtalker_temperatureoptFLOAT0.650.1–2—

Outputs (1)

NameTypeDescription
audioAUDIO—