Top TTS 2.5 - Synthesize
Make it say anything in someone else's voice, offline
- model
- reference_audio
- emotion
- emotion_audio
- audio
- generation_info
This is the whole reason you installed the pack. Top TTS 2.5 - Synthesize takes a few seconds of someone's voice, a line of text, and - if you want - an emotion, and returns speech in that voice with that feeling. Zero-shot voice cloning, five languages, fully local. No cloud, no uploads, no key. In the open TTS landscape this lands in the "quality + emotion + multilingual" corner: Chatterbox is the general default and F5-TTS the speed pick, but neither leans into multilingual and emotion control the way IndexTTS 2.5 does.
How it works
The node writes your reference audio to a temp WAV, then hands text and settings to the isolated worker process the Load Model node spawned. The worker seeds Python, NumPy, and PyTorch's RNGs from your seed, runs the vendored upstream IndexTTS pipeline - GPT-stage → s2mel → BigVGAN vocoder - and writes the result back out, which the node loads as a native ComfyUI AUDIO value. The model handle from the loader is required, so this node sits downstream of Top TTS 2.5 - Load Model in every workflow.
What a beginner actually sets
reference_audio- the voice you're cloning. Feed it a ComfyUIAUDIO(from a Load Audio node). The README's advice is the right advice: 5–15 seconds, single speaker, clean, low noise. And the community has learned this the hard way: loud, dynamic samples (podcast energy) clone far better than quiet, flat readings.text- the line to speak. The default is Chinese, so if you type English, set the language or you'll get... Chinese-pronounced English.language-ZH,EN,JA,ES, orAR. IndexTTS 2.5 is genuinely multilingual, not English-with-accents, which is a big part of its appeal.duration_factor- 1.0 is normal speed. Above 1 stretches (and slows generation), below 1 speeds up.seed- set it to reproduce a take; let it roll to explore.emotion_strength- how hard the emotion you've supplied hits, 0–1.
The optional emotion input is where the Emotion Vector node plugs in. You can also pass emotion_audio (a second clip carrying the feeling) or emotion_text (a phrase whose emotion gets read by the QwenEmotion model - but only if you enabled load_emotion_text_model in the loader, or you get an error). The pecking order is strict: vector or text outranks emotion_audio, and if you supply both a vector and emotion_text, the text-derived vector wins - the README claims vector-first, but the shipped code overwrites your vector with the text's. The rest - temperature, top_p, top_k, num_beams, repetition_penalty, max_mel_tokens, max_text_tokens_per_segment, interval_silence_ms, randomize_emotion, text_normalization - are the advanced drawer. Defaults are sane; you'll tweak repetition_penalty or temperature when output gets garbled or robotic, not before.
Outputs
audio (the AUDIO you wire into ComfyUI's core Save Audio / Preview Audio nodes) and generation_info, a string like IndexTTS 2.5 | EN | 3.42s | seed 12345 | local inference - handy if you're auto-naming files.
Troubleshooting
- English apostrophes break it. Real community finding: "don't" and "it's" get mangled. Write
dont,its. - Reference too quiet or flat → flat, poor clones. Re-record or pick a punchier clip before blaming the model.
- Emotion text fails → you forgot
load_emotion_text_model=Truein the loader. It's a clear error, but it's also more VRAM. - Worker died → check
ComfyUI/temp/comfyui_top_tts_worker.log; the ComfyUI console only shows the generic "worker error".
And the usual reminder that applies to every cloning tool: only clone voices you have the right to use, and check the upstream bilibili license before you ship anything commercial.
Inputs (20)
| Name | Type | Default | Description |
|---|---|---|---|
| model | TOP_TTS_2_5_MODEL | — | |
| text | STRING | 你好,欢迎使用 ComfyUI Top TTS。 | — |
| reference_audio | AUDIO | — | |
| language | COMBO | ZH | 5 options: ZH, EN, JA, ES, AR |
| duration_factor | FLOAT | 1.000.5–2 | — |
| emotion_strength | FLOAT | 1.000–1 | — |
| seed | INT | 00–4294967295 | — |
| emotionopt | TOP_TTS_2_5_EMOTION | — | |
| emotion_audioopt | AUDIO | — | |
| emotion_textopt | STRING | — | |
| randomize_emotionopt | BOOLEAN | false | — |
| interval_silence_msopt | INT | 2000–2000 | — |
| max_text_tokens_per_segmentopt | INT | 12020–600 | — |
| temperatureopt | FLOAT | 0.800.1–2 | — |
| top_popt | FLOAT | 0.800–1 | — |
| top_kopt | INT | 300–100 | — |
| num_beamsopt | INT | 31–10 | — |
| repetition_penaltyopt | FLOAT | 10.01–20 | — |
| max_mel_tokensopt | INT | 1500100–4000 | — |
| text_normalizationopt | BOOLEAN | true | — |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| audio | AUDIO | — |
| generation_info | STRING | — |