Nodes/ComfyUI-Lyonir-Studio/🐺 Lyonir Qwen3-TTS Voice Design
ComfyUI Node

🐺 Lyonir Qwen3-TTS Voice Design

Lyonir Qwen3-TTS Voice Design

By Lyonir·Created 4 days ago·Updated a day ago· 2
🐺 Lyonir Qwen3-TTS Voice Design
    • audio
    ◄textOlá! Esta é uma voz em português brasileiro.►
    ◄instructionHomem adulto, voz natural, cinematográfica, clara e expressiva. Fale de forma natural, expressiva e cinematográfica.►
    ◄model_choice1.7B►
    ◄deviceauto►
    ◄precisionbf16►
    ◄languagePortuguese (Brazil)►
    ◄seed0►
    ◄max_new_tokens2048►
    ◄top_p1.00►
    ◄top_k50►
    ◄temperature0.90►
    ◄repetition_penalty1.05►
    ◄attentionauto►
    ◄output_cleanupClean Voice (recommended)►
    ◄unload_model_after_generatefalse►
    ◄custom_model_path►
    ◄ptbr_engineNative PT-BR Hybrid (recommended)►
    ◄ptbr_checkpoint_step15000►
    ◄download_ptbr_if_missingtrue►

    Voice Design is the one Qwen3-TTS mode where you don't need a reference recording and you don't pick from a list of named speakers. You type a description - "adult man, natural, cinematic, clear and expressive" - and the model casts the voice. 🐺 Lyonir Qwen3-TTS Voice Design is the pack's wrapper around that, with the author's Brazilian-Portuguese obsession bolted on top.

    What it's for

    Audio in ComfyUI is the thinnest layer of the stack - a handful of bespoke packs with their own dependency trees, sitting off to the side of the checkpoint-and-sampler world. TTS in particular splits into three jobs, and one of them is casting. If you're voicing a character you don't have a recording of, and you can't be bothered to hunt a reference clip you're legally happy with, Voice Design is the shortest path: write the voice, generate the line.

    It's also the node to reach for when you're iterating on delivery rather than identity - angry, hushed, rushed, bored. Because there's no reference, nothing anchors the performance, so the instruction does all the work.

    How it works

    The backend is Qwen3-TTS running through qwen_tts, and Lyonir does not vendor it - it finds it. On first use the node scans sibling folders in custom_nodes for the backend that flybirdxx/ComfyUI-Qwen-TTS installs (it also accepts any folder containing qwen_tts/inference/qwen3_tts_model.py), adds that path, and patches a couple of mask functions so it works with your Transformers version. The deliberate choice here: it does not pip-install qwen-tts, because that package's metadata can drag a different Transformers into your environment. That's the dependency-hell story every ComfyUI user knows, avoided by hand.

    Then there's the PT-BR path, which is this pack's whole personality. For Brazilian Portuguese the node runs the pipeline in two stages: it generates an identity take from your instruction, then re-renders the target text through a dedicated Brazilian checkpoint as the accent and prosody source. The node description calls it "Brazil-first"; the tooltip for the engine choice is flat about it: all PT-BR modes use the dedicated Brazilian checkpoint as the accent source, none fall back to generic Portuguese. If you install nothing else, the checkpoint is snapshotted from Hugging Face into ComfyUI/models/qwen-tts/fala_pb_checkpoints/.

    Inputs and outputs

    Required: text (multiline, the line to speak), instruction (multiline, one unified field for identity, timbre, emotion, speed, energy and acting), model_choice - which is 1.7B and only 1.7B - plus device (auto/cuda/cpu/mps/xpu), precision (bf16 default) and language (12 options, defaulting to Portuguese (Brazil)).

    The instruction field is the node. Its tooltip: "Single instruction used for voice identity/design and delivery/performance. In PT-BR it is routed to both the identity stage and the Brazilian prosody/performance stage." Write it like a casting note, not a prompt.

    Optional but worth knowing: seed (same seed, same voice - that's your reproducibility lever, and the one to lock once you like a take), temperature (0.9 default; lower tightens delivery), output_cleanup with the tooltip "Post-synthesis cleanup for hiss/background noise", defaulting to Clean Voice (recommended), unload_model_after_generate if you share the GPU with a video model, attention (auto, or force sdpa/sage_attention/flash_attention_2/eager), max_new_tokens, top_p, top_k, repetition_penalty, and custom_model_path if you keep Qwen3-TTS weights somewhere non-standard.

    PT-BR-specific: ptbr_engine (4 modes, default "Native PT-BR Hybrid (recommended)"), ptbr_checkpoint_step (15000/10000/5000, default 15000), and download_ptbr_if_missing (on by default).

    Output: a single audio, typed AUDIO. Wire it into a Save Audio node, into Lyonir Save Video's audio input if you're scoring a clip, or into a lip-sync node downstream.

    Install

    The pack, plus the backend it borrows:

    cd ComfyUI/custom_nodes
    git clone https://github.com/Lyonir/ComfyUI-Lyonir-Studio.git
    git clone https://github.com/flybirdxx/ComfyUI-Qwen-TTS.git
    python -m pip install -r ComfyUI-Lyonir-Studio/requirements.txt
    

    Model weights live under ComfyUI/models/qwen-tts/. Restart ComfyUI, and hard-refresh the browser.

    Where people trip

    Qwen3-TTS backend not found. The node raises exactly that when it can't find qwen_tts anywhere - which happens if ComfyUI-Qwen-TTS isn't installed next to this pack in the same custom_nodes folder, or if its own dependencies failed on import. Check that pack's install first.

    First PT-BR run is slow and network-dependent. The Brazilian checkpoint downloads on demand. If you're offline or behind a proxy, turn download_ptbr_if_missing off and place the checkpoint yourself - otherwise you get a Hugging Face fetch error mid-generation.

    Generic Portuguese is not the same thing. If you type Brazilian text but leave language on the plain Portuguese option, you lose the whole dedicated path. Pick Portuguese (Brazil).

    Long text gets truncated mid-word. max_new_tokens caps generation at 2048 by default. Either raise it (in 256 steps, up to 8192) or split your script into lines - the second option is better practice anyway if you plan to edit the takes.

    Commercial use. The pack ships a NOTICE and a COMMERCIAL_LICENSES.md that the README tells you to read before deploying commercially. Voice models come with their own terms. Do that reading if you're selling something.

    CategoryLyonir Studio/Qwen3-TTS

    Inputs (19)

    NameTypeDefaultDescription
    textSTRINGOlá! Esta é uma voz em português brasileiro.—
    instructionSTRINGHomem adulto, voz natural, cinematográfica, clara e expressiva. Fale de forma natural, expressiva e cinematográfica.Single instruction used for voice identity/design and delivery/performance. In PT-BR it is routed to both the identity stage and the Brazilian prosody/performance stage.
    model_choiceCOMBO1.7B1 options: 1.7B
    deviceCOMBOauto5 options: auto, cuda, cpu, mps, xpu
    precisionCOMBObf163 options: bf16, fp16, fp32
    languageCOMBOPortuguese (Brazil)12 options: Auto, Chinese, English, Japanese, Korean, German, +6
    seedoptINT00–18446744073709550000—
    max_new_tokensoptINT2048256–8192—
    top_poptFLOAT1.000–1—
    top_koptINT500–200—
    temperatureoptFLOAT0.900.1–2—
    repetition_penaltyoptFLOAT1.051–2—
    attentionoptCOMBOauto5 options: auto, sage_attention, sdpa, eager, flash_attention_2
    output_cleanupoptCOMBOClean Voice (recommended)Post-synthesis cleanup for hiss/background noise. Strong Clean is more aggressive; Off returns the raw model output.
    unload_model_after_generateoptBOOLEANfalse—
    custom_model_pathoptSTRING—
    ptbr_engineoptCOMBONative PT-BR Hybrid (recommended)All PT-BR modes use the dedicated Brazilian checkpoint as the accent/prosody source.
    ptbr_checkpoint_stepoptCOMBO150003 options: 15000, 10000, 5000
    download_ptbr_if_missingoptBOOLEANtrue—

    Outputs (1)

    NameTypeDescription
    audioAUDIO—