Nodes/Steaked-nodes/Regional Prompts (Attention)
ComfyUI Node

Regional Prompts (Attention)

Four prompts, four boxes, one coherent image

By StealthNinja1O1·Created 11 months ago·Updated 5 months ago· 0
Regional Prompts (Attention)
  • clip
  • conditioning
◄base_prompt►
◄width1024►
◄height1024►
◄prompt_1►
◄prompt_2►
◄prompt_3►
◄prompt_4►

"Half of the image should be a castle, the other half a forest" is a request that drives prompt writers insane, because the model hears the whole sentence everywhere. Regional prompting is the fix: divide the canvas into boxes, give each box its own prompt, and let the sampler apply each prompt only inside its region. This node is the recommended flavor of that idea, and the one the pack's own README tells you to reach for first.

The trick is in the name. The Attention variant patches the model's cross-attention during sampling using what's essentially the "attention couple" technique that's been kicking around the community since early 2024 - it encodes each regional prompt separately, creates a mask for its box, and then, during every denoising step, biases the attention computation so each region's tokens only attend to their own pixels. The result is a single coherent image with genuinely different content per region, and - this is the selling point - no hard seams, because the blend is smooth at the boundaries. The tradeoff, per the author's own framing: some "bleeding" between regions, where one region's concepts leak into its neighbor.

What you actually set

  • base_prompt - the prompt that applies everywhere and fills any unmasked area. It's also prepended to each region's prompt, so it's your compositional backbone. Keep it general ("beautiful landscape, detailed, 8k") and let the boxes add specifics.
  • prompt_1 through prompt_4 - the per-region prompts. Empty ones are skipped.
  • width / height - canvas size (default 1024×1024), which must match your Empty Latent.
  • clip - your CLIP, straight from the checkpoint.

The node draws an interactive canvas on itself: drag to place up to four boxes, resize from the edges, and each selected box gets a weight (0–2, strength of that region's influence) plus start/end timestep controls (0–1). Setting start: 0.0, end: 0.3 applies a region's prompt only during the first 30% of sampling - great for stamping a layout early and letting detail generation take over. Box geometry persists in the workflow, so reloads keep your regions.

One conditioning output, wired into a KSampler in place of a normal positive encode.

Which variant to use

The pack ships six regional nodes. This one - plain text-to-image, no image input - is where you start. If you're starting from an existing image, you want the Img2Img or Image Input versions; if you're on an architecture where the attention-patch approach misbehaves, the Latent variant uses plain masked conditioning instead. For a first try, this one.

Install

Ships in Steaked-nodes:

cd ComfyUI/custom_nodes
git clone https://github.com/StealthNinja1O1/Steaked-nodes

Restart, or install via ComfyUI Manager ("Steaked-nodes"). Dependencies are torch/numpy/Pillow - nothing to download.

Common issues

If a region does nothing, check that its prompt isn't empty and its box is inside the canvas. If regions blend too much, drop the base prompt's descriptive power (it's prepended everywhere) or lower a region's weight. The attention hook path is powered by ComfyUI's hook system, so on older ComfyUI builds the node can silently fail to apply - update ComfyUI if you see it do nothing while your sampler runs normally.

CategorySteaked-nodes/prompting

Inputs (8)

NameTypeDefaultDescription
base_promptSTRING—
clipCLIP—
widthINT102464–8192—
heightINT102464–8192—
prompt_1optSTRING—
prompt_2optSTRING—
prompt_3optSTRING—
prompt_4optSTRING—

Outputs (1)

NameTypeDescription
conditioningCONDITIONING—