ComfyUI Node
LeapTalk Generate (portrait + speech)
512x512, 25 fps talking-head video with your audio, generated chunk by chunk by the LeapTalk runtime. Returns a VIDEO and a JSON report.
LeapTalk Generate (portrait + speech)
- runtime
- image
- audio
- video
- report
◄decoderlite_tae►
◄audio_guidance1.0►
◄previewtrue►
CategoryLeapTalk
Inputs (6)
| Name | Type | Default | Description |
|---|---|---|---|
| runtime | LEAPTALK_RUNTIME | — | |
| image | IMAGE | One reference portrait. The runtime resizes and center-crops it to 512x512, as the official script. | |
| audio | AUDIO | Speech to animate. It is muxed into the output unchanged; the model hears it resampled to 16 kHz mono. | |
| decoder | COMBO | lite_tae | lite_tae: LeapTalk's Lite decoder (taew2_1.pth, the official default). wan_vae: the standard Wan2.1 VAE from SoulX-FlashHead (slower; run without torch.compile here, the official script compiles it). |
| audio_guidance | FLOAT | 1.01–4 | Audio classifier-free guidance. 1.0 = off (the official default, 1 model call per chunk); >1 adds an unconditional call per chunk (2x model calls). Experimental: 2.0 gave visibly over-sharpened, discoloured lips in our tests; keep 1.0 unless you are experimenting. |
| preview | BOOLEAN | true | Show the last frame of each finished chunk while the job runs. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| video | VIDEO | — |
| report | STRING | — |