Nodes/ComfyUI Razv WaveSpeed Nodes/WaveSpeedAI InfiniteTalk
ComfyUI Node

WaveSpeedAI InfiniteTalk

A still photo that lip-syncs to your audio

By razvanmatei-sf·Created 12 months ago·Updated 9 months ago· 0
WaveSpeedAI InfiniteTalk
  • client
  • video_url
audio
image
resolution720p
enable_sync_modefalse
prompt
mask_image
seed-1
enable_base64_outputfalse

Take a photo of a person, hand it a piece of audio, and get back a video of that person saying it, with the mouth and expressions synced to the speech. That's InfiniteTalk - a talking-head / lip-sync model that hit the scene in August 2025 and got the internet very excited, very fast. This node from razvanmatei-sf/razv-wavespeed runs it through the WaveSpeed AI API, so no weights, no GPU, just an API key and a bill.

The obvious use is dubbing and avatars - voiceover a character, localize a video into another language while keeping the original performance's energy. The tooltip even mentions prompt can steer scene, pose, and behavior while preserving audio sync, so it's not a frozen-face puppet: you can direct some of the motion. There's also mask_image, to say which person in a group photo should do the talking. That last one is the detail that makes this node feel properly built rather than hastily wrapped.

How it works

WaveSpeedAI Clientimage (the face to animate) + audio (the speech track) → POST to the InfiniteTalk endpoint → poll → video_url out. The output is a hosted video URL, not a file - download it before it expires.

The inputs that actually matter:

  • audio - an audio file URL. The tooltip says connect from an "Upload Audio node," and just like the image-upload situation, that node isn't registered in the current build - so paste a publicly reachable audio URL.
  • image - the person to animate. Public URL.
  • resolution - 480p or 720p, and the tooltip prices it out: $0.15 per 5 seconds at 480p, $0.30 at 720p. Yes, the cost is in the node. That's this pack being honest.
  • prompt - optional instructions for scene/pose/behavior while keeping audio sync. This is the differentiator; use it.
  • mask_image - optional, to select which person in a multi-person image animates.
  • enable_base64_output - an optional toggle for getting the result as base64 instead of a URL. Leave it off unless you know why you need it.

seed (-1 = random) and enable_sync_mode (defaults off) round it out.

Install

Shared pack:

cd ComfyUI/custom_nodes
git clone https://github.com/razvanmatei-sf/razv-wavespeed
# restart ComfyUI

Or ComfyUI Manager → search "razv-wavespeed" → Install. Add a WaveSpeed AI key via the WaveSpeedAI Client node, config.ini (copy from config.ini.tmp, [API]), or a WAVESPEED_API_KEY env var.

Common issues

  • It's priced per 5 seconds and video-shaped, so a minute of output is real money. The tooltip tells you the exact cost; believe it.
  • URL inputs, and no upload node. The tooltips reference "Upload Audio"/"Upload Image" nodes that aren't in the current build. Host your files and paste the URLs.
  • Face quality sets the ceiling. A small, side-on, or low-res face won't lip-sync convincingly. Use a clean, frontal, well-lit headshot.
  • Sync fidelity varies with speech complexity - long silences, music, or multiple voices can confuse the alignment. Keep the audio track clean if you can.
  • Ethics note, because this class of model needs one: don't put words in a real person's mouth without consent. It's that simple.
CategoryWaveSpeedAI

Inputs (9)

NameTypeDefaultDescription
clientWAVESPEED_AI_API_CLIENT
audioSTRINGAudio file URL for generating lip-synced output (connect from Upload Audio node)
imageSTRINGImage to animate (connect from Upload Image node)
resolutionCOMBO720pOutput video resolution (480p: $0.15 per 5 seconds, 720p: $0.3 per 5 seconds)
enable_sync_modeBOOLEANfalseWait for video generation to complete before returning
promptoptSTRINGOptional generation instructions to control scene, pose, and behavior while maintaining audio synchronization
mask_imageoptSTRINGOptional mask image URL to specify which person to animate (connect from Upload Image node)
seedoptINT-1-1–2147483647Random seed for reproducible results. -1 for random seed
enable_base64_outputoptBOOLEANfalseEnable base64 output format

Outputs (1)

NameTypeDescription
video_urlSTRING