Kling Reference to Image
Subject, Scene, Style — Three Knobs For Character Consistency
- auth
- subject_image_1
- subject_image_2
- subject_image_3
- subject_image_4
- scene_image
- style_image
- image
- url
- task_id
This is the node that makes the consistency argument for going cloud-native. You hand Kling a picture of a subject, optionally a picture of a scene and a picture of a style, and it returns finished images that keep the subject while wearing the new setting. No LoRA, no training run, no ControlNet stack. Kling Reference to Image wraps Kling v2.1's multi-image2image endpoint, and it's the still-image sibling of the pack's multi-image video nodes.
Compare it honestly to what you can do locally. Qwen-Image-Edit plus a camera-angles LoRA gets you most of this, free and uncensored, and the community has largely settled on that as the default. What the local path still makes you earn is the seam: identity slides across a chain of edits, and the fix is crop-run-stitch. Kling charges per call and gives you nothing to maintain. If you're doing a one-off character sheet for a client, that's a good trade. If you're generating a hundred panels, it isn't.
How it works
Four types of reference are supported, and they are not interchangeable. subject_image_1 through subject_image_4 describe who or what - up to four angles of the same character or object. scene_image describes where. style_image describes how it's rendered. All of them get base64-encoded locally and posted to Kling's multi-image2image endpoint with model_name: kling-v2-1.
Then the node polls, which is the part beginners underestimate. It blocks the whole workflow until the job finishes, printing progress lines to the console, with a 1200-second timeout before it gives up. When the images land, it downloads every result and concatenates them into a single ComfyUI image batch, so generating four candidates gives you one image output with four frames in it.
The inputs worth caring about
subject_image_1is required; 2–4 are optional. Same warning as everywhere in this pack: crop to the subject first, because Kling does not crop. Full-frame shots with a tiny person in them produce results about the scenery.prompt- describe the change you want ("standing in a rain-soaked neon alley, medium shot"), not the person. The reference already does the person.aspect_ratio- eight choices, from16:9and21:9down to9:16for portrait work. Default is1:1.n- 1 to 9 images per call. Biggernmeans more to choose from and a bigger bill.scene_imageandstyle_imageare the two flexible slots. Use scene when the background has to be specific, style when you want the model to stop producing its own default look. Don't feed the same photo into both - you'll just confuse the conditioning.
Outputs are image (the batch of N generated images), url (the URL of the first image only, and it expires), and task_id. Because image is a batch, n: 9 and then a Save Image with prefix templating is the honest way to keep all nine.
Install
ComfyUI Manager → Install Custom Nodes → search "Kling Direct" → install → restart. Manual:
cd ComfyUI/custom_nodes
git clone https://github.com/IxMxAMAR/ComfyUI-Kling-Direct
Nothing else to fetch - the pack is stdlib plus requests/Pillow/numpy/torch/opencv-python, all already in ComfyUI. No checkpoints, no LoRAs, no model folder to fill.
You need keys from https://kling.ai/dev, which means clearing KYC once. Wire Kling AI Authentication → auth on this node. This node is fine with the access/secret pair; it's only Kling 3.0 Turbo that demands a separate API key. Not on the Singapore default region? Drop Kling Region Selector in between Auth and everything else, or your valid credentials will fail against the wrong gateway.
Things that go wrong
If the images come back looking like a different person entirely, nine times out of ten it's the crop. This node treats the reference as the whole subject, so a subject that occupies a small part of a landscape frame gets conditioned weakly.
If a task succeeds but the response has no usable image entries, the node throws rather than silently returning an empty batch - check the console for the response keys and try a smaller n.
And the mundane one: don't leave n at 9 while you're iterating on a prompt. Four candidates is usually enough to tell whether the composition works, and prompt iteration at n: 9 is how people discover the Cost Estimator node the hard way.
Inputs (10)
| Name | Type | Default | Description |
|---|---|---|---|
| auth | KLING_AUTH | — | |
| prompt | STRING | Text description of the image. | |
| subject_image_1 | IMAGE | Subject reference image. Crop to the subject first: Kling does not crop. | |
| aspect_ratio | COMBO | 1:1 | Output image aspect ratio. |
| n | INT | 11–9 | Number of images to generate (1-9). |
| subject_image_2opt | IMAGE | — | |
| subject_image_3opt | IMAGE | — | |
| subject_image_4opt | IMAGE | — | |
| scene_imageopt | IMAGE | Scene reference image. | |
| style_imageopt | IMAGE | Style reference image. |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| image | IMAGE | — |
| url | STRING | — |
| task_id | STRING | — |