Nodes/ComfyUI-Kling-Direct/Kling Reference to Image
ComfyUI Node

Kling Reference to Image

Subject, Scene, Style — Three Knobs For Character Consistency

By IxMxAMAR·Created 6 months ago·Updated 3 days ago· 5
Kling Reference to Image
  • auth
  • subject_image_1
  • subject_image_2
  • subject_image_3
  • subject_image_4
  • scene_image
  • style_image
  • image
  • url
  • task_id
◄prompt►
◄aspect_ratio1:1►
◄n1►

This is the node that makes the consistency argument for going cloud-native. You hand Kling a picture of a subject, optionally a picture of a scene and a picture of a style, and it returns finished images that keep the subject while wearing the new setting. No LoRA, no training run, no ControlNet stack. Kling Reference to Image wraps Kling v2.1's multi-image2image endpoint, and it's the still-image sibling of the pack's multi-image video nodes.

Compare it honestly to what you can do locally. Qwen-Image-Edit plus a camera-angles LoRA gets you most of this, free and uncensored, and the community has largely settled on that as the default. What the local path still makes you earn is the seam: identity slides across a chain of edits, and the fix is crop-run-stitch. Kling charges per call and gives you nothing to maintain. If you're doing a one-off character sheet for a client, that's a good trade. If you're generating a hundred panels, it isn't.

How it works

Four types of reference are supported, and they are not interchangeable. subject_image_1 through subject_image_4 describe who or what - up to four angles of the same character or object. scene_image describes where. style_image describes how it's rendered. All of them get base64-encoded locally and posted to Kling's multi-image2image endpoint with model_name: kling-v2-1.

Then the node polls, which is the part beginners underestimate. It blocks the whole workflow until the job finishes, printing progress lines to the console, with a 1200-second timeout before it gives up. When the images land, it downloads every result and concatenates them into a single ComfyUI image batch, so generating four candidates gives you one image output with four frames in it.

The inputs worth caring about

  • subject_image_1 is required; 2–4 are optional. Same warning as everywhere in this pack: crop to the subject first, because Kling does not crop. Full-frame shots with a tiny person in them produce results about the scenery.
  • prompt - describe the change you want ("standing in a rain-soaked neon alley, medium shot"), not the person. The reference already does the person.
  • aspect_ratio - eight choices, from 16:9 and 21:9 down to 9:16 for portrait work. Default is 1:1.
  • n - 1 to 9 images per call. Bigger n means more to choose from and a bigger bill.
  • scene_image and style_image are the two flexible slots. Use scene when the background has to be specific, style when you want the model to stop producing its own default look. Don't feed the same photo into both - you'll just confuse the conditioning.

Outputs are image (the batch of N generated images), url (the URL of the first image only, and it expires), and task_id. Because image is a batch, n: 9 and then a Save Image with prefix templating is the honest way to keep all nine.

Install

ComfyUI Manager → Install Custom Nodes → search "Kling Direct" → install → restart. Manual:

cd ComfyUI/custom_nodes
git clone https://github.com/IxMxAMAR/ComfyUI-Kling-Direct

Nothing else to fetch - the pack is stdlib plus requests/Pillow/numpy/torch/opencv-python, all already in ComfyUI. No checkpoints, no LoRAs, no model folder to fill.

You need keys from https://kling.ai/dev, which means clearing KYC once. Wire Kling AI Authentication → auth on this node. This node is fine with the access/secret pair; it's only Kling 3.0 Turbo that demands a separate API key. Not on the Singapore default region? Drop Kling Region Selector in between Auth and everything else, or your valid credentials will fail against the wrong gateway.

Things that go wrong

If the images come back looking like a different person entirely, nine times out of ten it's the crop. This node treats the reference as the whole subject, so a subject that occupies a small part of a landscape frame gets conditioned weakly.

If a task succeeds but the response has no usable image entries, the node throws rather than silently returning an empty batch - check the console for the response keys and try a smaller n.

And the mundane one: don't leave n at 9 while you're iterating on a prompt. Four candidates is usually enough to tell whether the composition works, and prompt iteration at n: 9 is how people discover the Cost Estimator node the hard way.

CategoryKling AI/Image

Inputs (10)

NameTypeDefaultDescription
authKLING_AUTH—
promptSTRINGText description of the image.
subject_image_1IMAGESubject reference image. Crop to the subject first: Kling does not crop.
aspect_ratioCOMBO1:1Output image aspect ratio.
nINT11–9Number of images to generate (1-9).
subject_image_2optIMAGE—
subject_image_3optIMAGE—
subject_image_4optIMAGE—
scene_imageoptIMAGEScene reference image.
style_imageoptIMAGEStyle reference image.

Outputs (3)

NameTypeDescription
imageIMAGE—
urlSTRING—
task_idSTRING—