Nodes/Steaked-nodes/Regional Prompts (Attention + Image Input)
ComfyUI Node

Regional Prompts (Attention + Image Input)

Regional prompting on anything upstream produces

By StealthNinja1O1·Created 11 months ago·Updated 5 months ago· 0
Regional Prompts (Attention + Image Input)
  • image
  • clip
  • image
  • conditioning
  • width
  • height
◄base_prompt►
◄prompt_1►
◄prompt_2►
◄prompt_3►
◄prompt_4►

The difference between "Img2Img" and "+ Image Input" in this pack's regional lineup is the difference between "pick a file" and "wire anything." The Img2Img variants take their image from ComfyUI's input-folder dropdown. These Image Input variants take a proper IMAGE socket, which means the source can be any node that outputs an image - a Load and Crop node, a preprocessor, a VAE decode, a saved generation - and the region canvas adapts to whatever resolution shows up.

Everything else is the attention-based regional machinery you've already seen: encode each region's prompt, mask it from its box, patch cross-attention during sampling so tokens stay in their box, and let the base prompt fill the rest. Outputs are the source image (passed through), the regional conditioning, and the source width/height - and this node is marked as an output node, so it can stand as a terminal in the graph.

Inputs and outputs

  • image - the IMAGE wire. Any upstream node that produces pixels.
  • base_prompt - everywhere prompt, also prepended to each region.
  • prompt_1 through prompt_4 - per-region prompts.
  • clip - your CLIP.

Outputs: image, conditioning, width, height.

Why the wire matters

Being fed over a wire changes the workflows you can build. The Img2Img variant is great when the source is a static file you chose. This one is for when the source changes: batch over a folder, pipe in a cropped region from Load and Crop, take the output of a previous generation and add regions on top, or run the same regional setup across different resolutions without touching a dropdown. The canvas reads dimensions from the incoming tensor, so if your source varies in size the boxes are laid out relative to whatever arrives - keep the source resolution consistent or your boxes will sit in the wrong place.

Install

Part of Steaked-nodes:

cd ComfyUI/custom_nodes
git clone https://github.com/StealthNinja1O1/Steaked-nodes

Restart, or ComfyUI Manager → "Steaked-nodes". No extra dependencies or downloads.

Common issues

The resolution-relative boxes are the main gotcha - a 512px source and a 1024px source put the same box geometry over totally different areas, so feed it a consistent size. As with all the img2img-family nodes, keep denoise strength moderate or the whole canvas re-rolls. And if a region seems inert, check its prompt isn't empty and its box is inside the incoming image; both are skipped silently. If you'd rather have hard boundaries than smooth blend, the Latent + Image Input sibling is the swap.

CategorySteaked-nodes/prompting

Inputs (7)

NameTypeDefaultDescription
imageIMAGE—
base_promptSTRING—
clipCLIP—
prompt_1optSTRING—
prompt_2optSTRING—
prompt_3optSTRING—
prompt_4optSTRING—

Outputs (4)

NameTypeDescription
imageIMAGE—
conditioningCONDITIONING—
widthINT—
heightINT—