Nodes/Comfyui-indextts25-xzg/IndexTTS 2.5 Speech Generation
ComfyUI Node

IndexTTS 2.5 Speech Generation

Multilingual zero-shot voice cloning with the IndexTTS 2.5 model, outputting standard ComfyUI AUDIO.

By xiaozhuguang·Created 9 days ago·Updated 5 days ago· 6
IndexTTS 2.5 Speech Generation
  • model
  • speaker_audio
  • emotion
  • sampling
  • Generated Audio
textWelcome to IndexTTS 2.5.
languageZH
duration_factor1.00
seed0
output_normalizationmatch reference
CategoryComfyui-indextts25-xzg/IndexTTS 2.5

Inputs (9)

NameTypeDefaultDescription
modelXZG_INDEXTTS25_MODEL
speaker_audioAUDIO
textSTRINGWelcome to IndexTTS 2.5.
languageCOMBOZH5 options: ZH, EN, JA, ES, AR
duration_factorFLOAT1.000.5–2
seedINT00–18446744073709550000
output_normalizationCOMBOmatch referencematch reference: the output is scaled to the same gated-RMS level as the speaker reference audio you fed in (loud in = loud out). rms -16 dB: fixed broadcast level with -1 dB peak ceiling. peak -1 dB: peak normalization only. off: raw model output.
emotionoptXZG_INDEXTTS25_EMOTIONFollows the voice reference when not connected.
samplingoptXZG_INDEXTTS25_SAMPLINGUses stable defaults when not connected.

Outputs (1)

NameTypeDescription
Generated AudioAUDIO