Top TTS 2.5 - Emotion Vector
Steering IndexTTS 2.5's emotion
- emotion
Most open TTS models give you one lever for feeling: a vague "expressiveness" dial that mostly makes the voice wobble. IndexTTS 2.5 is the odd one out, and this node is why. Top TTS 2.5 - Emotion Vector hands you eight separate emotion sliders - happy, angry, sad, afraid, disgusted, melancholic, surprised, calm - and turns them into a vector your cloned voice actually follows. It's the difference between a voice that says the words and a voice that sounds mad about them.
What it actually does
Mechanically it's embarrassingly simple: eight FLOAT sliders in, one TOP_TTS_2_5_EMOTION value out. The node just packs your slider values into an 8-dimensional list and passes it down the wire into the Synthesize node's optional emotion input. There's no model here, no inference, nothing to download. All the heavy lifting happens in Synthesize, where that vector gets mixed into the generation. Think of this as the knob box - the amp is downstream.
Two things are worth knowing before you start wiggling sliders:
- The vector and the text both outrank the emotion audio. Wire this node into Synthesize and your
emotion_audioreference gets ignored entirely - the node logic drops an emotion reference clip whenever a vector or emotion text is present, so it can't fight your dials. - The scale is 0–1.2, and it's not linear. A little goes a long way. The community experience with IndexTTS's emotion vectors has been that subtle values in the 0.2–0.3 range nudge delivery, while pushing several sliders toward 1.2 gives you full melodrama. You'll also notice the sliders don't all start at zero -
calmdefaults to 1.0, which is why out of the box everything sounds... calm.
The inputs that matter
You genuinely only need to touch two or three of these to get somewhere:
calm- the default 1.0 baseline. Drop it and raise one or two others to get emotion without everything else shifting.happy,angry,sad- the workhorses. Start at 0.2–0.3, listen, push.melancholic,surprised,afraid,disgusted- seasoning. Mostly useful once you know a voice well.
The emotion_strength input on the Synthesize node scales whatever vector you build here, so you can set a strong vector and fade it with one knob instead of retuning eight.
Install
This node ships in the ComfyUI-Top-TTS pack, so it comes along with the rest:
cd ComfyUI/custom_nodes
git clone https://github.com/whmc76/ComfyUI-Top-TTS.git
python -m pip install -r ComfyUI-Top-TTS/requirements.txt
python ComfyUI-Top-TTS/install.py
Then download the official weights (the pack refuses to run without them - see the Load Model article for the missing-files error):
cd ComfyUI-Top-TTS
python download_models.py --source huggingface --accept-license
You can also find it via ComfyUI Manager by searching "ComfyUI-Top-TTS".
Gotchas
Don't expect the vector to survive everything. If you're also passing emotion_text to Synthesize, that text path generates its own vector and takes priority over this node's output. And remember the vector steers delivery, not the voice itself - your reference speaker still supplies the timbre, so an angry vector on a soft, quiet reference reads different from an angry vector on a loud, punchy one. Give the reference the energy you want the output to have. The vector is a dial, not a switch.
Inputs (8)
| Name | Type | Default | Description |
|---|---|---|---|
| happy | FLOAT | 0.000–1.2 | — |
| angry | FLOAT | 0.000–1.2 | — |
| sad | FLOAT | 0.000–1.2 | — |
| afraid | FLOAT | 0.000–1.2 | — |
| disgusted | FLOAT | 0.000–1.2 | — |
| melancholic | FLOAT | 0.000–1.2 | — |
| surprised | FLOAT | 0.000–1.2 | — |
| calm | FLOAT | 1.000–1.2 | — |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| emotion | TOP_TTS_2_5_EMOTION | — |