Regional Prompts (Attention + Image Input)
Regional prompting on anything upstream produces
- image
- clip
- image
- conditioning
- width
- height
The difference between "Img2Img" and "+ Image Input" in this pack's regional lineup is the difference between "pick a file" and "wire anything." The Img2Img variants take their image from ComfyUI's input-folder dropdown. These Image Input variants take a proper IMAGE socket, which means the source can be any node that outputs an image - a Load and Crop node, a preprocessor, a VAE decode, a saved generation - and the region canvas adapts to whatever resolution shows up.
Everything else is the attention-based regional machinery you've already seen: encode each region's prompt, mask it from its box, patch cross-attention during sampling so tokens stay in their box, and let the base prompt fill the rest. Outputs are the source image (passed through), the regional conditioning, and the source width/height - and this node is marked as an output node, so it can stand as a terminal in the graph.
Inputs and outputs
image- the IMAGE wire. Any upstream node that produces pixels.base_prompt- everywhere prompt, also prepended to each region.prompt_1throughprompt_4- per-region prompts.clip- your CLIP.
Outputs: image, conditioning, width, height.
Why the wire matters
Being fed over a wire changes the workflows you can build. The Img2Img variant is great when the source is a static file you chose. This one is for when the source changes: batch over a folder, pipe in a cropped region from Load and Crop, take the output of a previous generation and add regions on top, or run the same regional setup across different resolutions without touching a dropdown. The canvas reads dimensions from the incoming tensor, so if your source varies in size the boxes are laid out relative to whatever arrives - keep the source resolution consistent or your boxes will sit in the wrong place.
Install
Part of Steaked-nodes:
cd ComfyUI/custom_nodes
git clone https://github.com/StealthNinja1O1/Steaked-nodes
Restart, or ComfyUI Manager → "Steaked-nodes". No extra dependencies or downloads.
Common issues
The resolution-relative boxes are the main gotcha - a 512px source and a 1024px source put the same box geometry over totally different areas, so feed it a consistent size. As with all the img2img-family nodes, keep denoise strength moderate or the whole canvas re-rolls. And if a region seems inert, check its prompt isn't empty and its box is inside the incoming image; both are skipped silently. If you'd rather have hard boundaries than smooth blend, the Latent + Image Input sibling is the swap.
Inputs (7)
| Name | Type | Default | Description |
|---|---|---|---|
| image | IMAGE | — | |
| base_prompt | STRING | — | |
| clip | CLIP | — | |
| prompt_1opt | STRING | — | |
| prompt_2opt | STRING | — | |
| prompt_3opt | STRING | — | |
| prompt_4opt | STRING | — |
Outputs (4)
| Name | Type | Description |
|---|---|---|
| image | IMAGE | — |
| conditioning | CONDITIONING | — |
| width | INT | — |
| height | INT | — |