Nodes/WavTTS/WavTTS Generate
ComfyUI Node

WavTTS Generate

Generate ComfyUI AUDIO from text using WavTTS zero-shot voice prompting.

By Saganaki22·Created 3 months ago·Updated 3 months ago· 7
WavTTS Generate
  • wavtts_model
  • reference_audio
  • audio
textHello! This is WavTTS running inside ComfyUI.
reference_text
Steps50
CFG3.0
speed1.00
timestep_mappingpower
timestep_power2.0
shift3.0
cross_fade_seconds0.00
fixed_total_duration_seconds0.0
seed0
CategoryWavTTS

Inputs (13)

NameTypeDefaultDescription
wavtts_modelWAVTTS_MODELLoaded WavTTS model from the WavTTS Load Model node.
reference_audioAUDIOReference speaker audio. Use a clean clip under 12 seconds.
textSTRINGHello! This is WavTTS running inside ComfyUI.Text to synthesize.
reference_textSTRINGTranscript of the reference audio. Required for WavTTS voice prompting.
StepsINT504–128Number of WavTTS sampling steps. Higher can improve quality but is slower.
CFGFLOAT3.00–10Classifier-free guidance strength. Higher values follow the text harder, but can sound less natural.
speedFLOAT1.000.1–3Speech pace multiplier used for duration estimation. 1.0 is normal.
timestep_mappingCOMBOpowerSampler timestep schedule. power is the upstream default; sway_sampling enables the EPS/Sway path.
timestep_powerFLOAT2.00.1–8Exponent for power timestep mapping. Only meaningful when timestep_mapping is power.
shiftFLOAT3.00.1–8Flow timestep shift. Lowering this can help if speech contains background noise.
cross_fade_secondsFLOAT0.000–2Overlap crossfade used only when long text is split into multiple generated chunks.
fixed_total_duration_secondsFLOAT0.00–1200 lets WavTTS estimate duration. Positive values force total prompt plus generated duration.
seedINT00–21474836470 is random/unseeded. Positive values make generation repeatable.

Outputs (1)

NameTypeDescription
audioAUDIO