Nodes/ComfyUI_Swwan/Image Crop By Mask And Resize (Swwan)
ComfyUI Node

Image Crop By Mask And Resize (Swwan)

The crop that hands you a bounding box to undo it with

By aining2022·Created 10 months ago·Updated about 18 hours ago· 33
Image Crop By Mask And Resize (Swwan)
  • image
  • mask
  • images
  • masks
  • bbox
◄base_resolution512►
◄padding0►
◄min_crop_resolution128►
◄max_crop_resolution512►

This is the node in the family that actually completes the detailing loop. Crop the masked region, upscale it to a size the model can render, sample it, and then paste it back exactly where it came from - and the paste-back needs coordinates, which is why this node has a bbox output.

If you've ever wondered why "only masked" workflows in ComfyUI tend to fall apart at the last step, it's usually this: the crop node didn't tell you where the crop went, so the composite had to guess.

Inputs and outputs

Required: image (IMAGE), mask (MASK), base_resolution (INT, default 512, step 8), padding (INT, default 0), min_crop_resolution (INT, default 128, step 8), max_crop_resolution (INT, default 512, step 8).

Outputs: images (IMAGE), masks (MASK), and bbox (BBOX).

The mechanism, which is not obvious

It processes the whole batch in two passes. First it finds a bounding box per frame: mask coordinates get rounded first (so ≥0.5 counts as inside), the box is then forced up to at least min_crop_resolution and down to at most max_crop_resolution on each axis, padding is added, and the box is clamped into the frame.

Then - this is the part that matters - it takes the maximum box across the entire batch, snaps that to a multiple of 16, and derives a single target size from base_resolution and the largest aspect ratio present. Every frame is then re-cropped around its own center using that shared maximum box, and all of them are resized to that same target.

Consequences:

  • All outputs are the same size. That's what makes them a legal tensor batch.
  • One large subject inflates every crop. A batch of five faces where one is huge means all five get the big box, so the small faces end up floating in extra context. If that bothers you, run the batch through Crop By Mask Batch instead, which sizes each crop independently.
  • Images are lanczos-resized, masks are bilinear. The mask goes soft at the edges, which is actually what you want for blending.
  • The target size is derived from base_resolution and the max aspect ratio, then rounded to /16. Setting base_resolution=1024 doesn't guarantee 1024-pixel output for every aspect ratio.

The bbox values are (x0, y0, x1, y1) per frame, in original-image coordinates - exactly what Image Uncrop By Mask consumes to put the re-rendered crop back down over the destination image. That's the intended pairing, and the README is explicit that BOX, BBOX, IMAGE_BOUNDS and SEAM are different protocols that can't be swapped: this is BBOX, and it goes into something that declares BBOX.

Where this fits

The classic use is the detailing loop the KB calls out - detect, crop and upscale, re-render, paste back - run once manually or per-frame across video. Because the crop lands at a usable resolution rather than 70 pixels of eye, the sampler gets a real generation budget for that region.

The second use is dataset preparation: crop a batch of masked subjects to uniform tiles for LoRA training, and keep the transforms so you can invert them.

Install

cd /path/to/ComfyUI/custom_nodes
git clone https://github.com/aining2022/ComfyUI_Swwan.git
cd ComfyUI_Swwan
python -m pip install -r requirements.txt

Windows portable:

.\python_embeded\python.exe -m pip install .\ComfyUI\custom_nodes\ComfyUI_Swwan\requirements.txt

Restart, hard-refresh, search Swwan; it sits under Swwan/Advanced/Image. No models, no GPUs required for the node itself. requirements.txt brings opencv-python, scipy and scikit-image; torch and torchvision come from your ComfyUI environment.

Common trip-ups, in order of how often they happen: setting max_crop_resolution below min_crop_resolution (the clamp order means min wins, silently); expecting per-frame, per-subject sizing from a mixed batch; and wiring bbox into something that wants BOX. That last one produces a confusing shape error, not a warning, so check the consumer's declared input type before you blame the crop.

CategorySwwan/Advanced/Image

Inputs (6)

NameTypeDefaultDescription
imageIMAGE—
maskMASK—
base_resolutionINT5120–16384—
paddingINT00–16384—
min_crop_resolutionINT1280–16384—
max_crop_resolutionINT5120–16384—

Outputs (3)

NameTypeDescription
imagesIMAGE—
masksMASK—
bboxBBOX—