ComfyUI Node
IndexTTS 2.5 Speech Generation
Multilingual zero-shot voice cloning with the IndexTTS 2.5 model, outputting standard ComfyUI AUDIO.
IndexTTS 2.5 Speech Generation
- model
- speaker_audio
- emotion
- sampling
- Generated Audio
◄textWelcome to IndexTTS 2.5.►
◄languageZH►
◄duration_factor1.00►
◄seed0►
◄output_normalizationmatch reference►
CategoryComfyui-indextts25-xzg/IndexTTS 2.5
Inputs (9)
| Name | Type | Default | Description |
|---|---|---|---|
| model | XZG_INDEXTTS25_MODEL | — | |
| speaker_audio | AUDIO | — | |
| text | STRING | Welcome to IndexTTS 2.5. | — |
| language | COMBO | ZH | 5 options: ZH, EN, JA, ES, AR |
| duration_factor | FLOAT | 1.000.5–2 | — |
| seed | INT | 00–18446744073709550000 | — |
| output_normalization | COMBO | match reference | match reference: the output is scaled to the same gated-RMS level as the speaker reference audio you fed in (loud in = loud out). rms -16 dB: fixed broadcast level with -1 dB peak ceiling. peak -1 dB: peak normalization only. off: raw model output. |
| emotionopt | XZG_INDEXTTS25_EMOTION | Follows the voice reference when not connected. | |
| samplingopt | XZG_INDEXTTS25_SAMPLING | Uses stable defaults when not connected. |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| Generated Audio | AUDIO | — |