🐺 Lyonir Qwen3-TTS Voice Clone
Clone a Voice From a Clip — and Let It Pick Its Own Best Take
- ref_audio
- audio
Voice cloning is the job that actually moved in the last couple of years - the KB's summary of the TTS field is that local closed the cloning gap, and Qwen3-TTS is one of the reasons. 🐺 Lyonir Qwen3-TTS Voice Clone is the reference-driven node of the pack: hand it a recording, hand it a line, get that voice saying your line. Then it quietly does something most clone nodes don't.
What it's for
Zero-shot cloning. No fine-tune, no training run, no waiting. You need a clean clip of the speaker - a few seconds of speech with no music, no room reverb, no overlapping voices - and the node does the rest. Wire its AUDIO output into the pack's video saver and you've got a character who can talk over the clip you just generated.
One expectation to set: cloning gives you identity, not acting. The upstream clone model is built around matching a speaker, and community reports of Qwen3-TTS in ComfyUI are clear that delivery control is not what the clone path is for. So treat the voice_instruction field on this node as exactly what the node description calls it - an "Instruction-First performance guide" that runs alongside a real reference supplying identity. It genuinely helps with pacing and emphasis. It is not a character slider.
How it works
The pack's clone module is explicit about its choices, and they're the right ones. It uses the ICL prompt path - reference audio plus an exact transcript - whenever it can, because that gives higher fidelity than feeding only the speaker embedding. Reference cleanup is deliberately conservative: mono conversion, DC removal, edge-silence trim, peak normalization. No denoising, no EQ, on purpose - those can alter the very timbre you're trying to clone.
Then there's candidate_count, which is the feature worth the download. Instead of generating one take and hoping, the node renders up to six candidates, embeds each one with Qwen's own speaker encoder, compares them by cosine similarity to the reference embedding (plus a lightweight delivery-similarity score against the guide audio), and returns the closest. Four candidates by default. It costs time; it buys consistency. If you're chasing a specific voice, that's a fair trade.
For Brazilian Portuguese there's a second layer, the same one the other two TTS nodes use: a dedicated Brazilian checkpoint acts as the accent and prosody source, so you never get generic Portuguese-accented output. All four brazilian_clone_mode options use it. That checkpoint gets snapshotted from Hugging Face into ComfyUI/models/qwen-tts/fala_pb_checkpoints/ when download_ptbr_if_missing is on.
Inputs and outputs
Required: ref_audio (AUDIO - your reference clip), target_text (multiline), model_choice (1.7B or 0.6B), device, precision, language (12 options). The model_choice tooltip notes that PT-BR Hybrid uses Base 1.7B internally for quality, while 0.6B applies to the official clone path for other languages - which is to say the PT-BR route ignores your 0.6B choice.
The optional pile is where the real knobs live: voice_instruction (the performance guide), brazilian_clone_mode, official_clone_mode - "Official Qwen ICL Clone (recommended)" or "Official Qwen Speaker-Only"; the tooltip explains ICL uses audio plus transcription for higher fidelity while Speaker-Only uses the vocal identity alone - ref_text, clone_profile (default Maximum Similarity), candidate_count (1–6, default 4), reference_processing (HQ Clean recommended, or Original), subtalker_temperature (0.65), plus the usual seed, temperature, top_p, top_k, repetition_penalty, max_new_tokens, attention, output_cleanup, unload_model_after_generate, custom_model_path, ptbr_checkpoint_step, download_ptbr_if_missing.
Output: one audio, typed AUDIO.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/Lyonir/ComfyUI-Lyonir-Studio.git
git clone https://github.com/flybirdxx/ComfyUI-Qwen-TTS.git
python -m pip install -r ComfyUI-Lyonir-Studio/requirements.txt
Weights go under ComfyUI/models/qwen-tts/. Restart ComfyUI and refresh the browser.
Where people trip
PT-BR clone demands a transcript. In Brazilian Portuguese mode, if ref_text is empty the node refuses outright - the error says PT-BR requires ref_text to be the exact transcript of the reference. That's not a suggestion; the accent anchor needs it. Type what the clip actually says, word for word, including "ums" if they're there.
Garbage in, garbage out. A clip with background music, a hissy mic, or two speakers will clone the noise along with the voice. Sixty seconds of clean, close-miked speech beats five minutes of podcast. Keep reference_processing on HQ Clean unless the reference is already pristine.
Cost sneaks up on you. candidate_count at 4 means four full generations plus embedding comparisons per run. Drop it to 1 while you're iterating on the text, then raise it for the take you keep.
Backend missing. If the node complains the Qwen3-TTS backend wasn't found, flybirdxx/ComfyUI-Qwen-TTS isn't installed alongside this pack in custom_nodes, or it failed to import. The pack borrows its bundled qwen_tts rather than pip-installing the standalone package, deliberately, to avoid a Transformers version fight.
Cloning someone and selling the result. Voice likeness is legally live territory - the pack ships NOTICE and COMMERCIAL_LICENSES.md and says to read them before commercial deployment. Read them.
Inputs (26)
| Name | Type | Default | Description |
|---|---|---|---|
| ref_audio | AUDIO | — | |
| target_text | STRING | Olá! Esta é uma clonagem de voz em português brasileiro. | — |
| model_choice | COMBO | 1.7B | PT-BR Hybrid usa Base 1.7B internamente para máxima qualidade. 0.6B é aplicado ao clone oficial dos outros idiomas. |
| device | COMBO | auto | 5 options: auto, cuda, cpu, mps, xpu |
| precision | COMBO | bf16 | 3 options: bf16, fp16, fp32 |
| language | COMBO | Portuguese (Brazil) | 12 options: Auto, Chinese, English, Japanese, Korean, German, +6 |
| seedopt | INT | 00–18446744073709550000 | — |
| max_new_tokensopt | INT | 2048256–8192 | — |
| top_popt | FLOAT | 1.000–1 | — |
| top_kopt | INT | 500–200 | — |
| temperatureopt | FLOAT | 0.900.1–2 | — |
| repetition_penaltyopt | FLOAT | 1.051–2 | — |
| attentionopt | COMBO | auto | 5 options: auto, sage_attention, sdpa, eager, flash_attention_2 |
| output_cleanupopt | COMBO | Clean Voice (recommended) | Post-synthesis cleanup for hiss/background noise. Strong Clean is more aggressive; Off returns the raw model output. |
| unload_model_after_generateopt | BOOLEAN | false | — |
| custom_model_pathopt | STRING | — | |
| voice_instructionopt | STRING | Descreva exatamente como a fala deve ser interpretada. Quando preenchida, o Lyonir usa uma rota Instruction-First para maximizar ritmo, emoção, energia, ênfase, pausas e entonação, preservando a identidade da referência. | — |
| brazilian_clone_modeopt | COMBO | Native PT-BR Hybrid (recommended) | Usado apenas quando language = Portuguese (Brazil). Todos os modos PT-BR usam o checkpoint brasileiro; nenhum usa o português genérico como fonte final de sotaque. |
| official_clone_modeopt | COMBO | Official Qwen ICL Clone (recommended) | Usado nos idiomas originais do Qwen. ICL usa áudio + transcrição para maior fidelidade; Speaker-Only usa apenas a identidade vocal. |
| ptbr_checkpoint_stepopt | COMBO | 15000 | Checkpoint 15000 é o recomendado para PT-BR. |
| download_ptbr_if_missingopt | BOOLEAN | true | — |
| ref_textopt | STRING | — | |
| clone_profileopt | COMBO | Maximum Similarity | Advanced setting. PT-BR Custom respeita este campo; os presets PT-BR recomendados usam configurações próprias. |
| candidate_countopt | INT | 41–6 | — |
| reference_processingopt | COMBO | HQ Clean (recommended) | 2 options: HQ Clean (recommended), Original |
| subtalker_temperatureopt | FLOAT | 0.650.1–2 | — |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| audio | AUDIO | — |