Nodes/LongCat AudioDiT TTS/LongCat AudioDiT Voice Clone TTS
ComfyUI Node

LongCat AudioDiT Voice Clone TTS

LongCat-AudioDiT Voice Clone TTS. Clones voice from reference audio using diffusion-based generation.

By Saganaki22·Created 4 months ago·Updated 4 months ago· 131
LongCat AudioDiT Voice Clone TTS
  • prompt_audio
  • audio
model_path
textThe sun glows warmly in a cloudless blue sky, a soft breeze drifts through the air, and birds fill the world with gentle, cheerful songs. Everything feels alive with beauty, just waiting to be discovered.
prompt_text
steps16
guidance_strength4.0
guidance_methodapg
deviceauto
dtypeauto
attentionauto
seed0
keep_model_loadedtrue
CategoryLongCat-AudioDiT

Inputs (12)

NameTypeDefaultDescription
model_pathCOMBOLongCat-AudioDiT model. Models are stored in ComfyUI/models/audiodit/
textSTRINGThe sun glows warmly in a cloudless blue sky, a soft breeze drifts through the air, and birds fill the world with gentle, cheerful songs. Everything feels alive with beauty, just waiting to be discovered.Text to synthesize in the cloned voice.
prompt_audioAUDIOReference audio to clone the voice from. 3-15 seconds gives the best results.
prompt_textSTRINGTranscript of the prompt audio. Required for voice cloning. Improves quality significantly.
stepsINT164–64Number of ODE Euler steps.
guidance_strengthFLOAT4.00–10CFG/APG guidance strength.
guidance_methodCOMBOapgGuidance method. 'apg' recommended for voice cloning.
deviceCOMBOautoCompute device.
dtypeCOMBOautoModel dtype.
attentionCOMBOautoAttention implementation.
seedINT00–2147483647Random seed. 0 = random.
keep_model_loadedBOOLEANtrueKeep model loaded between runs. Model is automatically offloaded to CPU after generation to free VRAM, then resumed to GPU on the next run.

Outputs (1)

NameTypeDescription
audioAUDIO