Nodes/ComfyUI_Swwan/CropByMask V3 (Swwan) · 旧版兼容
ComfyUI Node

CropByMask V3 (Swwan) · 旧版兼容

The mask-image variant, and its two odd quirks

By aining2022·Created 10 months ago·Updated about 18 hours ago· 33
CropByMask V3 (Swwan) · 旧版兼容
  • image
  • mask_image
  • crop_box
  • cropped_image
  • cropped_mask
  • crop_box
  • box_preview
◄detect▾►
◄top_reserve20►
◄bottom_reserve20►
◄left_reserve20►
◄right_reserve20►
◄round_to_multiple▾►

What's different about V3

Same family as CropByMask V2 - 孤海's LayerStyle node, re-registered by the Swwan pack as a legacy compatibility entry - but the interface changed in two ways that matter, and both of them are why people search for this node specifically.

First, V3 takes a mask the way the rest of ComfyUI does not: mask_image is an IMAGE, a black-and-white picture where white marks the region. It's not a MASK socket at all. If you've been feeding grayscale images around as pictures rather than masks, this is the node that speaks your dialect.

Second, both of its data inputs are optional. image, mask_image and crop_box are all optional; the node validates combinations instead. Give it only mask_image and it'll crop the mask itself. Give it only crop_box and a mask image, and it skips detection entirely.

The recommended replacement is CropByMask V5 (Swwan), and this one is here to preserve the historical data contract for old workflows. The old ID LayerUtility: CropByMask V3 isn't registered as an alias, so a workflow that references it won't resolve without migration.

How it works

The canvas size comes from image if you gave one, otherwise from mask_image. The mask image is converted to grayscale and, if crop_box wasn't supplied, blurred (radius 20) and measured with detect:

  • mask_area - bounding area of the white region.
  • min_bounding_rect - smallest enclosing rotated rectangle.
  • max_inscribed_rect - largest rectangle inside the region.

Then top_reserve, bottom_reserve, left_reserve, right_reserve (defaults 20, negatives allowed) push the box outward or pull it in, round_to_multiple (8 / 16 / 32 / 64 / 128 / 256 / 512 / None) snaps the crop to a friendly size and re-centres it, and the whole thing is clamped to the canvas.

Both rectangles get drawn on the preview: red for what detection found, green for the final crop. When detection is skipped because you passed a crop_box, only the green rectangle appears - which is a handy way to confirm your box is the one being used.

Inputs and outputs

Required: detect, the four reserves, round_to_multiple. Optional: image, mask_image, crop_box.

Outputs: cropped_image, cropped_mask, crop_box, box_preview.

That second output is the quirk worth knowing: cropped_mask is an IMAGE, not a MASK. V2 returns a MASK there; V3 returns the cropped mask rendered as a picture. Wire it into a MASK socket and it won't connect, and the error will look like a type mismatch rather than a version difference. This trips up people reading V2 docs while running V3, which is basically everybody.

Also: if you provide image, the image gets cropped and the mask output is the cropped mask. If you don't provide an image at all, cropped_image becomes the cropped mask rendered as RGB - so a mask-only call still gives you something displayable.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/aining2022/ComfyUI_Swwan
cd ComfyUI_Swwan
python -m pip install -r requirements.txt

Manager users: search ComfyUI Swwan. Restart and hard-refresh afterwards. No models, no optional pip extras - this node rides on the base requirements (numpy, Pillow, opencv-python, scipy, scikit-image).

Errors and traps

"Either mask_image or crop_box must be provided" is the validation firing because you gave it neither. "At least one of image or mask_image must be provided" means you gave it a crop box and nothing to crop. Both are honest errors; the first is usually someone wiring a MASK into mask_image and getting a conversion that ends up empty.

Only the first frame is measured. The mask image's first frame decides the geometry, so a batch of different-sized or differently-positioned frames gets one crop.

Passing crop_box forward is the real workflow. Run V3 once to detect, then reuse the BOX output downstream to crop a second image with identical geometry - that's how you keep a detailer crop and its paste-back in lockstep. The box is a plain (x1, y1, x2, y2) sequence, and it is not interchangeable with BBOX or IMAGE_BOUNDS from elsewhere in the pack; the protocols differ.

And the naming smell that catches everyone eventually: ComfyUI Manager's "conflict with ComfyUI_Swwan" note is about node names overlapping other packs' - LayerStyle in this case. It's a warning about coexistence, not a broken install, and the current pack registers its own IDs rather than claiming the old aliases.

CategorySwwan/Legacy

Inputs (9)

NameTypeDefaultDescription
detectCOMBO3 options: mask_area, min_bounding_rect, max_inscribed_rect
top_reserveINT20-9999–9999—
bottom_reserveINT20-9999–9999—
left_reserveINT20-9999–9999—
right_reserveINT20-9999–9999—
round_to_multipleCOMBO8 options: 8, 16, 32, 64, 128, 256, +2
imageoptIMAGE—
mask_imageoptIMAGE—
crop_boxoptBOX—

Outputs (4)

NameTypeDescription
cropped_imageIMAGE—
cropped_maskIMAGE—
crop_boxBOX—
box_previewIMAGE—