Nodes/ComfyUI_AIIA/IndexTTS-2 TTS
ComfyUI Node

IndexTTS-2 TTS

A ComfyUI node in AIIA/Synthesis with 20 inputs and 1 output.

By havvk·Created about a year ago·Updated 6 months ago· 12
IndexTTS-2 TTS
  • indextts_model
  • reference_audio
  • emotion_audio
  • audio
text你好,这是一段 IndexTTS-2 语音合成测试。
voice_presetFemale_HQ
emo_alpha1.00
happy0.00
angry0.00
sad0.00
afraid0.00
disgusted0.00
melancholic0.00
surprised0.00
calm0.00
use_emo_textfalse
emo_text
interval_silence200
max_text_tokens_per_segment120
use_randomfalse
seed0
CategoryAIIA/Synthesis

Inputs (20)

NameTypeDefaultDescription
indextts_modelINDEXTTS_MODEL
textSTRING你好,这是一段 IndexTTS-2 语音合成测试。
voice_presetCOMBOFemale_HQ4 options: Female_HQ, Male_HQ, Female, Male
reference_audiooptAUDIOSpeaker voice reference. Leave empty to use voice_preset.
emotion_audiooptAUDIOOptional emotion reference audio (separate from speaker voice).
emo_alphaoptFLOAT1.000–1Emotion blending strength (0=no emotion, 1=full emotion).
happyoptFLOAT0.000–1
angryoptFLOAT0.000–1
sadoptFLOAT0.000–1
afraidoptFLOAT0.000–1
disgustedoptFLOAT0.000–1
melancholicoptFLOAT0.000–1
surprisedoptFLOAT0.000–1
calmoptFLOAT0.000–1
use_emo_textoptBOOLEANfalseAuto-detect emotion from text using built-in Qwen emotion model. Overrides emotion sliders.
emo_textoptSTRINGCustom emotion text prompt (used with use_emo_text). Leave empty to use main text.
interval_silenceoptINT2000–2000Silence duration (ms) inserted between text segments for long text.
max_text_tokens_per_segmentoptINT12030–500Max tokens per text segment. Lower = more segments, higher = longer per-segment generation.
use_randomoptBOOLEANfalseEnable random sampling (reduces voice cloning fidelity).
seedoptINT0-1–2147483647Random seed. -1 = random.

Outputs (1)

NameTypeDescription
audioAUDIO