Kling Multi-Image to Video
Four Photos of the Same Subject, One Coherent Clip
- auth
- image_1
- image_2
- image_3
- image_4
- video
- video_file
- audio
- url
- task_id
Normal image-to-video takes one picture and animates it. Kling Multi-Image to Video takes up to four and uses them as subject references - the model is told what the character looks like from several angles, then generates new footage of that character rather than just nudging your first frame. It's the difference between "make this photo move" and "make a clip starring this person."
This is a Kling v1.6 feature, which matters if you're comparing it to the newer stuff. Kling 2.5/3.0 and the Omni models are where the image quality went. Multi-Image is the cheap workhorse you reach for when what you actually need is the right face, and you're willing to accept older-model fidelity to get it.
The mechanism
The node base64-encodes each connected image, posts them as an image_list to Kling's multi-image2video endpoint with model_name: kling-v1-6, and polls the task until it's done. When the task finishes, it downloads the finished MP4 into ComfyUI's output folder under a random kling_<hex>.mp4 name, then decodes that file twice: once through OpenCV into a frame batch for the video output, and once through torchaudio into the audio output.
Worth knowing: the wait is real. This node blocks while polling, up to a 1200-second ceiling (20 minutes), printing elapsed seconds to your console as it goes. You can cancel a running job and ComfyUI will pick that up within about a second - but the request itself is already paid for by then.
Inputs that matter
The required set is prompt, image_1, negative_prompt, mode, duration, aspect_ratio; image_2 through image_4 are optional and you should fill them if you have them.
image_1- the primary subject reference. The tooltip carries the single most important warning in this article: crop to the subject first, because Kling does not crop. A full-body shot where the person occupies a tenth of the frame gives you a clip about a background. Crop to head-and-shoulders or tighter.mode-proorstd. Std first, always. You're test-running a composition, not mastering a film, and pro will happily double your bill for a take you're going to reject.duration- the strings"5"or"10", not an integer. Nine seconds is not an option here.aspect_ratio-16:9,9:16,1:1.negative_prompt- plainer than you're used to from Stable Diffusion. "blurry, extra limbs, distorted face" is plenty; the 40-token SDXL negative stack does nothing useful on a closed video model.
Outputs: video (IMAGE batch - frames), video_file (the filename in your output dir), audio, url (the temporary Kling CDN link, which expires), and task_id. Send video to Video Combine or a save node; send audio to Fast Video Saver if you want the clip with sound; keep task_id if you want to re-check the job without re-running it.
Install
Manager → search "Kling Direct" → install → restart. Manual:
cd ComfyUI/custom_nodes
git clone https://github.com/IxMxAMAR/ComfyUI-Kling-Direct
No model downloads, no extra Python packages beyond what ComfyUI ships. Credentials come from https://kling.ai/dev - access key plus secret key, KYC required on a fresh account. Add the Kling AI Authentication node and wire auth in. If your account is not in the Singapore region, put a Kling Region Selector between Auth and this node, otherwise every call authenticates fine and then 401s on the endpoint.
Common problems
Images that come back as solid frames of nothing, or a "download failed" error, usually mean the payload was rejected upstream. This pack falls back from PNG to JPEG when an image blows past Kling's ~10 MB request ceiling, and that re-encode is lossy - if you're feeding it a 4K plate, resize first instead of letting the node quietly crush it.
Faces that drift mid-clip are the classic complaint about every reference-image video model, and it's why the tooltip nags you about cropping. Tighter crop in, better face out.
And a warning on cost: Kling V3 Omni runs about $1.50 for a five-second 1080p clip. This node is the v1.6 path and cheaper, but you're still metering every take. Run Cost Estimator before you send fifty variations into the queue - the subscription-lock objection people have about API nodes is mostly a surprise bill objection, and it's avoidable.
Inputs (10)
| Name | Type | Default | Description |
|---|---|---|---|
| auth | KLING_AUTH | — | |
| prompt | STRING | Text description of the video. | |
| image_1 | IMAGE | Subject reference image. Crop to the subject first: Kling does not crop. | |
| negative_prompt | STRING | Things to avoid in the generated video. | |
| mode | COMBO | std | Generation mode: 'pro' for higher quality, 'std' for faster/cheaper. |
| duration | COMBO | 5 | Video duration in seconds. |
| aspect_ratio | COMBO | 16:9 | Output video aspect ratio. |
| image_2opt | IMAGE | — | |
| image_3opt | IMAGE | — | |
| image_4opt | IMAGE | — |
Outputs (5)
| Name | Type | Description |
|---|---|---|
| video | IMAGE | — |
| video_file | STRING | — |
| audio | AUDIO | — |
| url | STRING | — |
| task_id | STRING | — |