Kling Avatar 2.0
A broadcast-ready talking avatar from one photo and an audio file
- image
- sound_file
- VIDEO
Kling Avatar 2.0 is the node that turns a single photo into a broadcast-style digital human. One reference image, one audio file, and Kling generates a video of that person talking, moving, and emoting like a news anchor or a product presenter. It's the step up from plain lip sync: lip sync (Kling Lip Sync Video with Audio) reanimates a video of a face, while Avatar starts from a still photo and builds the whole performance from scratch - mouth, expressions, and head motion included. Both are built-in Kling partner nodes, meaning cloud API calls billed through your Comfy account, no local model.
Where this earns its keep is exactly where local talking-head tools (LivePortrait and friends) get fiddly: it's a one-click "photo + audio = presenter" pipeline with production-ready framing, which is why the "Avatar" name is doing real work here. The classic use is digital presenters - training explainer videos, product pages, dubbing a talking head into another language - where you control the person's likeness with a photo and the script with audio.
The inputs
- image - the avatar's reference photo. Constraints from the tooltip: width and height at least 300px, and an aspect ratio between 1:2.5 and 2.5:1 (so no extreme panoramic or ultra-tall crops).
- sound_file - the audio that drives the performance. Must be between 2 and 300 seconds - under 2s is too short to do anything with; over 300s is rejected.
- mode -
stdorpro. Pro is the higher-quality, higher-cost tier; std is the draft tier. Budget accordingly. - seed - re-run control only; results are nondeterministic regardless of seed.
- prompt (optional) - the tooltip is the pitch: define the avatar's actions, emotions, and camera movements. Without it you get a neutral read of the script; with it you get direction - "look up and smile at the end," "gesture while explaining."
What comes out
A single VIDEO output - your photo, animated into a speaking performance.
Gotchas
- The photo is the ceiling. A low-res or heavily angled reference photo limits how "broadcast" the result looks. Use a clear, front-facing headshot at the best resolution you have.
- Prompt is optional but powerful. A blank prompt gives a flat presentation; the node's reason to exist is the direction you add. If your avatars look robotic, that's the knob you forgot.
- Audio quality gates the sync. Background noise and reverb hurt the lip-sync, same as any talking-head model. Clean audio in, clean presenter out.
- Per-second pricing, pro-tier more so. A 5-minute avatar in pro mode is a real spend. Draft in std, then render the keep in pro.
- Watch the seed disclaimer. It only decides whether the node re-runs; you can't reproduce a specific take by replaying a seed.
If you've ever needed "a person saying your words" without renting a studio or a face, this is the node. Photo, audio, a prompt for direction, and a presenter comes out the other side.
Inputs (5)
| Name | Type | Default | Description |
|---|---|---|---|
| image | IMAGE | Avatar reference image. Width and height must be at least 300px. Aspect ratio must be between 1:2.5 and 2.5:1. | |
| sound_file | AUDIO | Audio input. Must be between 2 and 300 seconds in duration. | |
| mode | COMBO | 2 options: std, pro | |
| seed | INT | 00–2147483647 | Seed controls whether the node should re-run; results are non-deterministic regardless of seed. |
| promptopt | STRING | Optional prompt to define avatar actions, emotions, and camera movements. |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| VIDEO | VIDEO | — |