Nodes/ComfyUI/HeyGen Talking Photo
ComfyUI Node Runs on cloud

HeyGen Talking Photo

Make any portrait talk

By Comfy-Org·Created 4 years ago·Updated about 5 hours ago· 128,055
HeyGen Talking Photo
  • image
  • VIDEO
speech
resolution1080p
aspect_ratioauto
expressivenesslow
seed42

The photo is in your prompt history, or your downloads folder, or a past workflow you can't quite find. This node takes any image of a person and animates it into a talking video - real lip sync, natural head motion, the works - driven by a script or your own audio. You've probably seen the results everywhere; this is the node that makes them in ComfyUI.

It's HeyGen's Avatar IV engine doing the animation on HeyGen's servers, billed per second through your Comfy account. It landed in ComfyUI core in July 2026. The workflow that keeps me coming back: generate a portrait locally with SDXL or Flux - any style, any face - and hand it here for the talking. That combo, local still + cloud animation, is what makes this node feel like cheating.

How it works

The image is uploaded (downscaled automatically if it's bigger than 2K), then the node calls HeyGen's video API with the image as the anchor frame. The speech input decides how it moves: script runs HeyGen text-to-speech - pick a voice, optionally a custom voice ID, tweak speed - and audio lip-syncs your own recording (up to 10 minutes). Either way the mouth follows the speech, and expressiveness dials how much the face and hands get into it.

The inputs that matter

  • image - the portrait. Single clear face, good lighting; the whole result hinges on this.
  • speech - script (text + voice) or audio. For a TTS-only chain, HeyGen Text to Speech feeds naturally into the audio side of this.
  • expressiveness - low, medium, high. Low is the safe default: professional and controlled. High risks the uncanny valley on a low-quality source.
  • resolution, aspect_ratio - output size; "auto" aspect follows the input image, which keeps your composition intact.
  • seed - a dummy; not sent to HeyGen, just there to force a re-run.

Output is one VIDEO.

Gotchas

The quality ceiling is your input. A mid-shot, front-facing, well-lit face gives the best animation; profile or heavily angled shots degrade fast, and weird crops make the model improvise. expressiveness is the other trap - everyone cranks it to high and gets a puppet, then walks it back to low. And remember the meter: it's priced per second of output, so a long script is a long bill. Keep scripts tight, preview with a short take, and you've got a legitimate talking head with zero filming.

Categorypartner/video/HeyGen

Inputs (6)

NameTypeDefaultDescription
imageIMAGEImage of a person to animate. Downscaled automatically if larger than 2K.
speechCOMBODrive the avatar with a text script (HeyGen text-to-speech) or your own audio.
resolutionoptCOMBO1080pOutput video resolution.
aspect_ratiooptCOMBOautoOutput aspect ratio. 'auto' follows the input image.
expressivenessoptCOMBOlowHow expressive the animated face and gestures are.
seedoptINT420–2147483647Not sent to HeyGen; change it to force a re-run.

Outputs (1)

NameTypeDescription
VIDEOVIDEO