Nodes/ComfyUI-Lyonir-Studio/๐Ÿบ Lyonir Qwen3-TTS Custom Voice
ComfyUI Node

๐Ÿบ Lyonir Qwen3-TTS Custom Voice

Ten Named Voices, One Instruction Field โ€” Lyonir Qwen3-TTS Custom Voice

By LyonirยทCreated 4 days agoยทUpdated a day agoยท 2
๐Ÿบ Lyonir Qwen3-TTS Custom Voice
    • audio
    โ—„textOlรก! Esta รฉ uma voz brasileira nativa.โ–บ
    โ—„speakerLeonardoโ–บ
    โ—„model_choice1.7Bโ–บ
    โ—„deviceautoโ–บ
    โ—„precisionbf16โ–บ
    โ—„languagePortuguese (Brazil)โ–บ
    โ—„seed0โ–บ
    โ—„max_new_tokens2048โ–บ
    โ—„top_p1.00โ–บ
    โ—„top_k50โ–บ
    โ—„temperature0.90โ–บ
    โ—„repetition_penalty1.05โ–บ
    โ—„attentionautoโ–บ
    โ—„output_cleanupClean Voice (recommended)โ–บ
    โ—„unload_model_after_generatefalseโ–บ
    โ—„custom_model_pathโ–บ
    โ—„instructionโ–บ
    โ—„ptbr_engineNative PT-BR Hybrid (recommended)โ–บ
    โ—„ptbr_checkpoint_step15000โ–บ
    โ—„download_ptbr_if_missingtrueโ–บ
    โ—„custom_speaker_nameโ–บ

    There are three ways to get a voice out of Qwen3-TTS in this pack. Voice Design invents one from a description. Voice Clone copies one from a recording. Custom Voice sits in the middle: you pick from a roster of ten built-in speakers, and then you direct them.

    What it's for

    Sometimes you don't want a new voice, you want the same voice across forty lines and thirteen episodes, with the acting changing line to line. A named speaker gives you that stability for free - same name, same voice, every run. On top of that, one instruction field lets you steer emotion, pacing, energy and delivery without changing the casting.

    That makes this the node for episodic work: narration, character dialogue in a series, an audiobook, anything with continuity. The KB's read on local TTS applies here - the tooling is real and good, but it lives in bespoke packs with their own dependency stacks rather than in ComfyUI's center. This node is one of those packs.

    The ten speakers, and the Leonardo quirk

    speaker is a combo of ten choices: Leonardo, Aiden, Dylan, Eric, Ono_Anna, Ryan, Serena, Sohee, Uncle_Fu, Vivian. Default is Leonardo - the pack's own default, unsurprising from an author whose whole node set is Brazil-first.

    Here's the thing you need to know before you spend twenty minutes wondering why your voice keeps coming out male: the node description says "Leonardo is always forced to an adult male Aiden-based identity." It's not a bug and it's not subtle in the code either - the instruction is prefixed with an explicit masculinity directive, and the terminal prints that it's enforcing an adult male Aiden identity. If you were hoping to steer Leonardo somewhere feminine with the instruction field, you can't. Pick another speaker.

    Inputs and how the pipeline works

    Required: text (multiline), speaker, model_choice (1.7B or 0.6B), device, precision (bf16 default), language (12 options).

    Then there's instruction, a single multiline box that's optional here - unlike Voice Design, where it's a required input. Its tooltip: "Single instruction used for voice identity/design and delivery/performance. In PT-BR it is routed to both the identity stage and the Brazilian prosody/performance stage." In other words it isn't a one-pass prompt hint. For Brazilian Portuguese the node generates an identity take, then re-renders through a dedicated Brazilian checkpoint as the accent and prosody source - one instruction, two stages.

    That's also where the downloads happen. ptbr_engine (4 modes, default "Native PT-BR Hybrid (recommended)") and ptbr_checkpoint_step (15000/10000/5000, 15000 being the recommended one) select the Brazilian checkpoint; download_ptbr_if_missing (on) fetches it from Hugging Face into ComfyUI/models/qwen-tts/fala_pb_checkpoints/ on first use.

    Sampling controls, all optional and all defaulted sensibly: seed, temperature (0.9), top_p (1.0), top_k (50), repetition_penalty (1.05), max_new_tokens (2048, in 256 steps up to 8192), attention (auto or force a backend), output_cleanup - the author's own words: "Post-synthesis cleanup for hiss/background noise. Strong Clean is more aggressive; Off returns the raw model output", defaulting to Clean Voice (recommended) - unload_model_after_generate for GPU sharing, custom_model_path if your weights aren't under models/qwen-tts/, and custom_speaker_name for a custom Qwen CustomVoice model with a speaker id that isn't in the ten.

    Output: one audio, typed AUDIO. Into a Save Audio node, or straight into Lyonir Save Video's audio input to score a clip.

    How it runs under the hood

    The node doesn't ship Qwen3-TTS. It locates the backend installed by flybirdxx/ComfyUI-Qwen-TTS in a sibling folder under custom_nodes, adds it to the path, and patches compatibility around your Transformers version. It deliberately won't pip-install the qwen-tts package - that package's metadata can pin a different Transformers and wreck the environment you already have. Respectable instinct.

    Install

    cd ComfyUI/custom_nodes
    git clone https://github.com/Lyonir/ComfyUI-Lyonir-Studio.git
    git clone https://github.com/flybirdxx/ComfyUI-Qwen-TTS.git
    python -m pip install -r ComfyUI-Lyonir-Studio/requirements.txt
    

    Then restart ComfyUI and refresh the browser.

    Where people trip

    Backend not found. The node's error message is explicit: keep ComfyUI-Qwen-TTS installed because it provides the vendored qwen_tts backend, or supply an existing qwen_tts package without changing your Transformers version. If that pack failed to import, this node fails too.

    1.7B versus 0.6B. 1.7B is the default and the quality one; 0.6B is the one you reach for when VRAM is tight. Same speakers, smaller model, less polish.

    Instruction field left empty. You still get a voice - but you've thrown away the acting control, which is half the reason to use this node instead of a plain TTS node.

    Language left on generic Portuguese for Brazilian text. You lose the dedicated Brazilian accent path entirely. Set language to Portuguese (Brazil).

    Repeat runs eat VRAM. Each run loads a model. unload_model_after_generate is off by default because reloading is slow; turn it on if you're sharing the GPU with a video sampler.

    Commercial deployment. The README points at NOTICE and COMMERCIAL_LICENSES.md before you sell anything made with these voice nodes. Read them.

    CategoryLyonir Studio/Qwen3-TTS

    Inputs (21)

    NameTypeDefaultDescription
    textSTRINGOlรก! Esta รฉ uma voz brasileira nativa.โ€”
    speakerCOMBOLeonardo10 options: Leonardo, Aiden, Dylan, Eric, Ono_Anna, Ryan, +4
    model_choiceCOMBO1.7B2 options: 1.7B, 0.6B
    deviceCOMBOauto5 options: auto, cuda, cpu, mps, xpu
    precisionCOMBObf163 options: bf16, fp16, fp32
    languageCOMBOPortuguese (Brazil)12 options: Auto, Chinese, English, Japanese, Korean, German, +6
    seedoptINT00โ€“18446744073709550000โ€”
    max_new_tokensoptINT2048256โ€“8192โ€”
    top_poptFLOAT1.000โ€“1โ€”
    top_koptINT500โ€“200โ€”
    temperatureoptFLOAT0.900.1โ€“2โ€”
    repetition_penaltyoptFLOAT1.051โ€“2โ€”
    attentionoptCOMBOauto5 options: auto, sage_attention, sdpa, eager, flash_attention_2
    output_cleanupoptCOMBOClean Voice (recommended)Post-synthesis cleanup for hiss/background noise. Strong Clean is more aggressive; Off returns the raw model output.
    unload_model_after_generateoptBOOLEANfalseโ€”
    custom_model_pathoptSTRINGโ€”
    instructionoptSTRINGSingle instruction used for voice identity/design and delivery/performance. In PT-BR it is routed to both the identity stage and the Brazilian prosody/performance stage.
    ptbr_engineoptCOMBONative PT-BR Hybrid (recommended)All PT-BR modes use the dedicated Brazilian checkpoint as the accent/prosody source.
    ptbr_checkpoint_stepoptCOMBO150003 options: 15000, 10000, 5000
    download_ptbr_if_missingoptBOOLEANtrueโ€”
    custom_speaker_nameoptSTRINGโ€”

    Outputs (1)

    NameTypeDescription
    audioAUDIOโ€”