Image Crop By Mask Batch (Swwan)
One image, many masks, uniform tiles out
- image
- masks
- images
- masks
Detector-driven pipelines produce a bag of masks for a single image - twelve faces, four hands, whatever the model found. You want twelve crops that are all the same size so they can be a batch, all sampled in one pass, then pasted back where they came from.
That's this node. It's the batch-side companion to Image Crop By Mask And Resize, and the difference between them matters: that one takes a batch and produces uniform crops using a shared maximum box; this one takes one image and produces one crop per mask, each sized to fit your target independently.
Inputs and outputs
Required: image (IMAGE), masks (MASK), width (INT, default 512, step 8), height (INT, default 512, step 8), padding (INT, default 0, 0–4096), preserve_size (BOOLEAN, default false), and bg_color (STRING, default 0, 0, 0, tooltip: "Color as RGB values in range 0-255, separated by commas").
Outputs: images (IMAGE) and masks (MASK) - matching batches, so the crop and its mask stay aligned through whatever you do next.
How it works
For each mask in the batch: find the non-zero bounding box, expand it by padding, cut that rectangle out of the image, then fit it into a width × height canvas - scaled down with lanczos to preserve aspect if it's too big, or (when preserve_size is on and the crop is already smaller than the target) left at native size - and centered. Finally, everything outside the mask inside that tile gets flattened to bg_color.
Two details worth flagging because they're easy to misread:
- Only the first image is used. The code crops
image[0]for every mask, so the image input is effectively a single frame. Masks are the batch dimension here, not images. If you feed a batch of frames expecting per-frame crops, you'll get every crop from frame one. - Masks are resized to the image if they disagree. If your detector ran on a different resolution than your image, the mask batch is nearest-resized to match - no error, just a silent reconciliation you should be aware of.
There's also a hole-punching behavior: masks with nothing in them are continued, so they simply don't produce output. That means output index N doesn't map back to mask index N if any mask was empty. If you're recombining, keep the masks output and match on it rather than on position.
And if no mask produces anything, you get empty tensors of the right shape - (0, height, width, 3) - not an exception. A downstream node that doesn't handle a zero-length batch will be the thing that complains.
Where it fits
The canonical pipeline: a detector emits SEGS or a mask batch → this node gives you uniform crops of the subject(s) → those go through an upscale and an img2img/detail pass at a real resolution → the results go back for compositing. The KB's framing of the whole category applies: an automated detailer is just "zoom in on a mask, upscale, re-diffuse," and every detailer differs only in where the mask came from.
The other use is extraction: a subject-per-mask contact sheet where every tile is the same convenient square, ready for a dataset folder or a grid comparison.
If you want to compose these back onto the original, padding is the knob that makes the seams behave - a little padding gives the composite some context to blend into, which is exactly what a bare bbox crop denies you.
Install
cd /path/to/ComfyUI/custom_nodes
git clone https://github.com/aining2022/ComfyUI_Swwan.git
cd ComfyUI_Swwan
python -m pip install -r requirements.txt
Windows portable:
.\python_embeded\python.exe -m pip install .\ComfyUI\custom_nodes\ComfyUI_Swwan\requirements.txt
Restart ComfyUI, hard-refresh the browser, search Swwan. It's filed under Swwan/Advanced/Batch and needs no detection models of its own - you supply the masks. Dependencies are whatever requirements.txt installs (opencv-python, scipy, scikit-image) over your existing torch and Pillow, Python ≥3.11.
One install-flavored gotcha, since this is a mask node and mask nodes get blamed for everything: if your crops come out inverted - background kept, subject gone - check where the mask came from. ComfyUI's Load Image emits 1 - alpha for the MASK output, so a hand-drawn mask derived from a Load Image alpha is upside down relative to what this node wants. The pack's own RGBA notes call this out explicitly; it's the single most common "this node is broken" report in masking workflows.
Inputs (7)
| Name | Type | Default | Description |
|---|---|---|---|
| image | IMAGE | — | |
| masks | MASK | — | |
| width | INT | 5120–16384 | — |
| height | INT | 5120–16384 | — |
| padding | INT | 00–4096 | — |
| preserve_size | BOOLEAN | false | — |
| bg_color | STRING | 0, 0, 0 | Color as RGB values in range 0-255, separated by commas. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| images | IMAGE | — |
| masks | MASK | — |