🐺 Lyonir Qwen3-TTS Voice Design
Lyonir Qwen3-TTS Voice Design
- audio
Voice Design is the one Qwen3-TTS mode where you don't need a reference recording and you don't pick from a list of named speakers. You type a description - "adult man, natural, cinematic, clear and expressive" - and the model casts the voice. 🐺 Lyonir Qwen3-TTS Voice Design is the pack's wrapper around that, with the author's Brazilian-Portuguese obsession bolted on top.
What it's for
Audio in ComfyUI is the thinnest layer of the stack - a handful of bespoke packs with their own dependency trees, sitting off to the side of the checkpoint-and-sampler world. TTS in particular splits into three jobs, and one of them is casting. If you're voicing a character you don't have a recording of, and you can't be bothered to hunt a reference clip you're legally happy with, Voice Design is the shortest path: write the voice, generate the line.
It's also the node to reach for when you're iterating on delivery rather than identity - angry, hushed, rushed, bored. Because there's no reference, nothing anchors the performance, so the instruction does all the work.
How it works
The backend is Qwen3-TTS running through qwen_tts, and Lyonir does not vendor it - it finds it. On first use the node scans sibling folders in custom_nodes for the backend that flybirdxx/ComfyUI-Qwen-TTS installs (it also accepts any folder containing qwen_tts/inference/qwen3_tts_model.py), adds that path, and patches a couple of mask functions so it works with your Transformers version. The deliberate choice here: it does not pip-install qwen-tts, because that package's metadata can drag a different Transformers into your environment. That's the dependency-hell story every ComfyUI user knows, avoided by hand.
Then there's the PT-BR path, which is this pack's whole personality. For Brazilian Portuguese the node runs the pipeline in two stages: it generates an identity take from your instruction, then re-renders the target text through a dedicated Brazilian checkpoint as the accent and prosody source. The node description calls it "Brazil-first"; the tooltip for the engine choice is flat about it: all PT-BR modes use the dedicated Brazilian checkpoint as the accent source, none fall back to generic Portuguese. If you install nothing else, the checkpoint is snapshotted from Hugging Face into ComfyUI/models/qwen-tts/fala_pb_checkpoints/.
Inputs and outputs
Required: text (multiline, the line to speak), instruction (multiline, one unified field for identity, timbre, emotion, speed, energy and acting), model_choice - which is 1.7B and only 1.7B - plus device (auto/cuda/cpu/mps/xpu), precision (bf16 default) and language (12 options, defaulting to Portuguese (Brazil)).
The instruction field is the node. Its tooltip: "Single instruction used for voice identity/design and delivery/performance. In PT-BR it is routed to both the identity stage and the Brazilian prosody/performance stage." Write it like a casting note, not a prompt.
Optional but worth knowing: seed (same seed, same voice - that's your reproducibility lever, and the one to lock once you like a take), temperature (0.9 default; lower tightens delivery), output_cleanup with the tooltip "Post-synthesis cleanup for hiss/background noise", defaulting to Clean Voice (recommended), unload_model_after_generate if you share the GPU with a video model, attention (auto, or force sdpa/sage_attention/flash_attention_2/eager), max_new_tokens, top_p, top_k, repetition_penalty, and custom_model_path if you keep Qwen3-TTS weights somewhere non-standard.
PT-BR-specific: ptbr_engine (4 modes, default "Native PT-BR Hybrid (recommended)"), ptbr_checkpoint_step (15000/10000/5000, default 15000), and download_ptbr_if_missing (on by default).
Output: a single audio, typed AUDIO. Wire it into a Save Audio node, into Lyonir Save Video's audio input if you're scoring a clip, or into a lip-sync node downstream.
Install
The pack, plus the backend it borrows:
cd ComfyUI/custom_nodes
git clone https://github.com/Lyonir/ComfyUI-Lyonir-Studio.git
git clone https://github.com/flybirdxx/ComfyUI-Qwen-TTS.git
python -m pip install -r ComfyUI-Lyonir-Studio/requirements.txt
Model weights live under ComfyUI/models/qwen-tts/. Restart ComfyUI, and hard-refresh the browser.
Where people trip
Qwen3-TTS backend not found. The node raises exactly that when it can't find qwen_tts anywhere - which happens if ComfyUI-Qwen-TTS isn't installed next to this pack in the same custom_nodes folder, or if its own dependencies failed on import. Check that pack's install first.
First PT-BR run is slow and network-dependent. The Brazilian checkpoint downloads on demand. If you're offline or behind a proxy, turn download_ptbr_if_missing off and place the checkpoint yourself - otherwise you get a Hugging Face fetch error mid-generation.
Generic Portuguese is not the same thing. If you type Brazilian text but leave language on the plain Portuguese option, you lose the whole dedicated path. Pick Portuguese (Brazil).
Long text gets truncated mid-word. max_new_tokens caps generation at 2048 by default. Either raise it (in 256 steps, up to 8192) or split your script into lines - the second option is better practice anyway if you plan to edit the takes.
Commercial use. The pack ships a NOTICE and a COMMERCIAL_LICENSES.md that the README tells you to read before deploying commercially. Voice models come with their own terms. Do that reading if you're selling something.
Inputs (19)
| Name | Type | Default | Description |
|---|---|---|---|
| text | STRING | Olá! Esta é uma voz em português brasileiro. | — |
| instruction | STRING | Homem adulto, voz natural, cinematográfica, clara e expressiva. Fale de forma natural, expressiva e cinematográfica. | Single instruction used for voice identity/design and delivery/performance. In PT-BR it is routed to both the identity stage and the Brazilian prosody/performance stage. |
| model_choice | COMBO | 1.7B | 1 options: 1.7B |
| device | COMBO | auto | 5 options: auto, cuda, cpu, mps, xpu |
| precision | COMBO | bf16 | 3 options: bf16, fp16, fp32 |
| language | COMBO | Portuguese (Brazil) | 12 options: Auto, Chinese, English, Japanese, Korean, German, +6 |
| seedopt | INT | 00–18446744073709550000 | — |
| max_new_tokensopt | INT | 2048256–8192 | — |
| top_popt | FLOAT | 1.000–1 | — |
| top_kopt | INT | 500–200 | — |
| temperatureopt | FLOAT | 0.900.1–2 | — |
| repetition_penaltyopt | FLOAT | 1.051–2 | — |
| attentionopt | COMBO | auto | 5 options: auto, sage_attention, sdpa, eager, flash_attention_2 |
| output_cleanupopt | COMBO | Clean Voice (recommended) | Post-synthesis cleanup for hiss/background noise. Strong Clean is more aggressive; Off returns the raw model output. |
| unload_model_after_generateopt | BOOLEAN | false | — |
| custom_model_pathopt | STRING | — | |
| ptbr_engineopt | COMBO | Native PT-BR Hybrid (recommended) | All PT-BR modes use the dedicated Brazilian checkpoint as the accent/prosody source. |
| ptbr_checkpoint_stepopt | COMBO | 15000 | 3 options: 15000, 10000, 5000 |
| download_ptbr_if_missingopt | BOOLEAN | true | — |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| audio | AUDIO | — |