Nodes/ComfyUI-LeapTalk/LeapTalk Generate (portrait + speech)
ComfyUI Node

LeapTalk Generate (portrait + speech)

512x512, 25 fps talking-head video with your audio, generated chunk by chunk by the LeapTalk runtime. Returns a VIDEO and a JSON report.

By hiroki-abe-58·Created 2 days ago·Updated 2 days ago· 1
LeapTalk Generate (portrait + speech)
  • runtime
  • image
  • audio
  • video
  • report
◄decoderlite_tae►
◄audio_guidance1.0►
◄previewtrue►
CategoryLeapTalk

Inputs (6)

NameTypeDefaultDescription
runtimeLEAPTALK_RUNTIME—
imageIMAGEOne reference portrait. The runtime resizes and center-crops it to 512x512, as the official script.
audioAUDIOSpeech to animate. It is muxed into the output unchanged; the model hears it resampled to 16 kHz mono.
decoderCOMBOlite_taelite_tae: LeapTalk's Lite decoder (taew2_1.pth, the official default). wan_vae: the standard Wan2.1 VAE from SoulX-FlashHead (slower; run without torch.compile here, the official script compiles it).
audio_guidanceFLOAT1.01–4Audio classifier-free guidance. 1.0 = off (the official default, 1 model call per chunk); >1 adds an unconditional call per chunk (2x model calls). Experimental: 2.0 gave visibly over-sharpened, discoloured lips in our tests; keep 1.0 unless you are experimenting.
previewBOOLEANtrueShow the last frame of each finished chunk while the job runs.

Outputs (2)

NameTypeDescription
videoVIDEO—
reportSTRING—