๐บ Lyonir Qwen3-TTS Custom Voice
Ten Named Voices, One Instruction Field โ Lyonir Qwen3-TTS Custom Voice
- audio
There are three ways to get a voice out of Qwen3-TTS in this pack. Voice Design invents one from a description. Voice Clone copies one from a recording. Custom Voice sits in the middle: you pick from a roster of ten built-in speakers, and then you direct them.
What it's for
Sometimes you don't want a new voice, you want the same voice across forty lines and thirteen episodes, with the acting changing line to line. A named speaker gives you that stability for free - same name, same voice, every run. On top of that, one instruction field lets you steer emotion, pacing, energy and delivery without changing the casting.
That makes this the node for episodic work: narration, character dialogue in a series, an audiobook, anything with continuity. The KB's read on local TTS applies here - the tooling is real and good, but it lives in bespoke packs with their own dependency stacks rather than in ComfyUI's center. This node is one of those packs.
The ten speakers, and the Leonardo quirk
speaker is a combo of ten choices: Leonardo, Aiden, Dylan, Eric, Ono_Anna, Ryan, Serena, Sohee, Uncle_Fu, Vivian. Default is Leonardo - the pack's own default, unsurprising from an author whose whole node set is Brazil-first.
Here's the thing you need to know before you spend twenty minutes wondering why your voice keeps coming out male: the node description says "Leonardo is always forced to an adult male Aiden-based identity." It's not a bug and it's not subtle in the code either - the instruction is prefixed with an explicit masculinity directive, and the terminal prints that it's enforcing an adult male Aiden identity. If you were hoping to steer Leonardo somewhere feminine with the instruction field, you can't. Pick another speaker.
Inputs and how the pipeline works
Required: text (multiline), speaker, model_choice (1.7B or 0.6B), device, precision (bf16 default), language (12 options).
Then there's instruction, a single multiline box that's optional here - unlike Voice Design, where it's a required input. Its tooltip: "Single instruction used for voice identity/design and delivery/performance. In PT-BR it is routed to both the identity stage and the Brazilian prosody/performance stage." In other words it isn't a one-pass prompt hint. For Brazilian Portuguese the node generates an identity take, then re-renders through a dedicated Brazilian checkpoint as the accent and prosody source - one instruction, two stages.
That's also where the downloads happen. ptbr_engine (4 modes, default "Native PT-BR Hybrid (recommended)") and ptbr_checkpoint_step (15000/10000/5000, 15000 being the recommended one) select the Brazilian checkpoint; download_ptbr_if_missing (on) fetches it from Hugging Face into ComfyUI/models/qwen-tts/fala_pb_checkpoints/ on first use.
Sampling controls, all optional and all defaulted sensibly: seed, temperature (0.9), top_p (1.0), top_k (50), repetition_penalty (1.05), max_new_tokens (2048, in 256 steps up to 8192), attention (auto or force a backend), output_cleanup - the author's own words: "Post-synthesis cleanup for hiss/background noise. Strong Clean is more aggressive; Off returns the raw model output", defaulting to Clean Voice (recommended) - unload_model_after_generate for GPU sharing, custom_model_path if your weights aren't under models/qwen-tts/, and custom_speaker_name for a custom Qwen CustomVoice model with a speaker id that isn't in the ten.
Output: one audio, typed AUDIO. Into a Save Audio node, or straight into Lyonir Save Video's audio input to score a clip.
How it runs under the hood
The node doesn't ship Qwen3-TTS. It locates the backend installed by flybirdxx/ComfyUI-Qwen-TTS in a sibling folder under custom_nodes, adds it to the path, and patches compatibility around your Transformers version. It deliberately won't pip-install the qwen-tts package - that package's metadata can pin a different Transformers and wreck the environment you already have. Respectable instinct.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/Lyonir/ComfyUI-Lyonir-Studio.git
git clone https://github.com/flybirdxx/ComfyUI-Qwen-TTS.git
python -m pip install -r ComfyUI-Lyonir-Studio/requirements.txt
Then restart ComfyUI and refresh the browser.
Where people trip
Backend not found. The node's error message is explicit: keep ComfyUI-Qwen-TTS installed because it provides the vendored qwen_tts backend, or supply an existing qwen_tts package without changing your Transformers version. If that pack failed to import, this node fails too.
1.7B versus 0.6B. 1.7B is the default and the quality one; 0.6B is the one you reach for when VRAM is tight. Same speakers, smaller model, less polish.
Instruction field left empty. You still get a voice - but you've thrown away the acting control, which is half the reason to use this node instead of a plain TTS node.
Language left on generic Portuguese for Brazilian text. You lose the dedicated Brazilian accent path entirely. Set language to Portuguese (Brazil).
Repeat runs eat VRAM. Each run loads a model. unload_model_after_generate is off by default because reloading is slow; turn it on if you're sharing the GPU with a video sampler.
Commercial deployment. The README points at NOTICE and COMMERCIAL_LICENSES.md before you sell anything made with these voice nodes. Read them.
Inputs (21)
| Name | Type | Default | Description |
|---|---|---|---|
| text | STRING | Olรก! Esta รฉ uma voz brasileira nativa. | โ |
| speaker | COMBO | Leonardo | 10 options: Leonardo, Aiden, Dylan, Eric, Ono_Anna, Ryan, +4 |
| model_choice | COMBO | 1.7B | 2 options: 1.7B, 0.6B |
| device | COMBO | auto | 5 options: auto, cuda, cpu, mps, xpu |
| precision | COMBO | bf16 | 3 options: bf16, fp16, fp32 |
| language | COMBO | Portuguese (Brazil) | 12 options: Auto, Chinese, English, Japanese, Korean, German, +6 |
| seedopt | INT | 00โ18446744073709550000 | โ |
| max_new_tokensopt | INT | 2048256โ8192 | โ |
| top_popt | FLOAT | 1.000โ1 | โ |
| top_kopt | INT | 500โ200 | โ |
| temperatureopt | FLOAT | 0.900.1โ2 | โ |
| repetition_penaltyopt | FLOAT | 1.051โ2 | โ |
| attentionopt | COMBO | auto | 5 options: auto, sage_attention, sdpa, eager, flash_attention_2 |
| output_cleanupopt | COMBO | Clean Voice (recommended) | Post-synthesis cleanup for hiss/background noise. Strong Clean is more aggressive; Off returns the raw model output. |
| unload_model_after_generateopt | BOOLEAN | false | โ |
| custom_model_pathopt | STRING | โ | |
| instructionopt | STRING | Single instruction used for voice identity/design and delivery/performance. In PT-BR it is routed to both the identity stage and the Brazilian prosody/performance stage. | |
| ptbr_engineopt | COMBO | Native PT-BR Hybrid (recommended) | All PT-BR modes use the dedicated Brazilian checkpoint as the accent/prosody source. |
| ptbr_checkpoint_stepopt | COMBO | 15000 | 3 options: 15000, 10000, 5000 |
| download_ptbr_if_missingopt | BOOLEAN | true | โ |
| custom_speaker_nameopt | STRING | โ |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| audio | AUDIO | โ |