Nodes/ComfyUI/Kling Avatar 2.0
ComfyUI Node Runs on cloud

Kling Avatar 2.0

A broadcast-ready talking avatar from one photo and an audio file

By Comfy-Org·Created 4 years ago·Updated a day ago· 130,493
Kling Avatar 2.0
  • image
  • sound_file
  • VIDEO
mode
seed0
prompt

Kling Avatar 2.0 is the node that turns a single photo into a broadcast-style digital human. One reference image, one audio file, and Kling generates a video of that person talking, moving, and emoting like a news anchor or a product presenter. It's the step up from plain lip sync: lip sync (Kling Lip Sync Video with Audio) reanimates a video of a face, while Avatar starts from a still photo and builds the whole performance from scratch - mouth, expressions, and head motion included. Both are built-in Kling partner nodes, meaning cloud API calls billed through your Comfy account, no local model.

Where this earns its keep is exactly where local talking-head tools (LivePortrait and friends) get fiddly: it's a one-click "photo + audio = presenter" pipeline with production-ready framing, which is why the "Avatar" name is doing real work here. The classic use is digital presenters - training explainer videos, product pages, dubbing a talking head into another language - where you control the person's likeness with a photo and the script with audio.

The inputs

  • image - the avatar's reference photo. Constraints from the tooltip: width and height at least 300px, and an aspect ratio between 1:2.5 and 2.5:1 (so no extreme panoramic or ultra-tall crops).
  • sound_file - the audio that drives the performance. Must be between 2 and 300 seconds - under 2s is too short to do anything with; over 300s is rejected.
  • mode - std or pro. Pro is the higher-quality, higher-cost tier; std is the draft tier. Budget accordingly.
  • seed - re-run control only; results are nondeterministic regardless of seed.
  • prompt (optional) - the tooltip is the pitch: define the avatar's actions, emotions, and camera movements. Without it you get a neutral read of the script; with it you get direction - "look up and smile at the end," "gesture while explaining."

What comes out

A single VIDEO output - your photo, animated into a speaking performance.

Gotchas

  • The photo is the ceiling. A low-res or heavily angled reference photo limits how "broadcast" the result looks. Use a clear, front-facing headshot at the best resolution you have.
  • Prompt is optional but powerful. A blank prompt gives a flat presentation; the node's reason to exist is the direction you add. If your avatars look robotic, that's the knob you forgot.
  • Audio quality gates the sync. Background noise and reverb hurt the lip-sync, same as any talking-head model. Clean audio in, clean presenter out.
  • Per-second pricing, pro-tier more so. A 5-minute avatar in pro mode is a real spend. Draft in std, then render the keep in pro.
  • Watch the seed disclaimer. It only decides whether the node re-runs; you can't reproduce a specific take by replaying a seed.

If you've ever needed "a person saying your words" without renting a studio or a face, this is the node. Photo, audio, a prompt for direction, and a presenter comes out the other side.

Categorypartner/video/Kling

Inputs (5)

NameTypeDefaultDescription
imageIMAGEAvatar reference image. Width and height must be at least 300px. Aspect ratio must be between 1:2.5 and 2.5:1.
sound_fileAUDIOAudio input. Must be between 2 and 300 seconds in duration.
modeCOMBO2 options: std, pro
seedINT00–2147483647Seed controls whether the node should re-run; results are non-deterministic regardless of seed.
promptoptSTRINGOptional prompt to define avatar actions, emotions, and camera movements.

Outputs (1)

NameTypeDescription
VIDEOVIDEO