Image Crop By Mask (Swwan)
Autocrop to whatever your mask is pointing at
- image
- mask
- image
Every detailing pipeline is the same four steps: detect a region, crop it, re-render it at a resolution the model can actually resolve, paste it back. This node is step two for the simple case - you already have a mask, and you want just the part of the image the mask covers.
It's the plainest member of the family: two inputs, one output, no options. That minimalism is the feature and also the limitation.
Inputs and outputs
Required: image (IMAGE) and mask (MASK). Output: image.
The mask is rounded first - so in practice anything ≥0.5 counts as inside, anything below is outside - and then the bounding box of the non-zero region is computed per frame in the batch. The image gets cut to that box, and every frame in the batch is returned as a batch.
Two behaviors worth knowing:
- Masks are matched to frames by index, and if the mask batch is shorter, the last mask is reused for the remaining frames. That's a convenience when you have one mask for a whole batch, and a silent surprise when you have fewer masks than frames by accident.
- An empty mask passes that frame through uncropped. No error, no black frame - the original pixels. If you're filtering that later by size, don't assume every output frame is a crop.
The sharp edge
If the frames in your batch crop to different sizes, this node raises: Mask crops have different sizes; use Mask Crop or process images as a list. It doesn't pad, and it doesn't center. That's the author being deliberate about a rule the README states too - a tensor batch has to be uniform, so a heterogeneous set of crops can't be one.
If you're processing a batch of detected faces, that's the case you'll hit immediately, and it's exactly why the pack also ships SwwanImageCropByMaskBatch (fixed output size, per-mask crops from one image) and SwwanImageCropByMaskAndResize (uniform target size, plus a bbox output). Those two are what you want for any real detector-driven pipeline.
When to reach for this one
Single-image autocrop before an img2img pass. Detect a subject, crop to it, run the crop through a sampler at a resolution that's actually useful, then composite back. For the composite step read the output contract carefully: this node emits no bounding box, so Image Uncrop By Mask (which requires a BBOX) has nothing to work with from here. Use Image Crop By Mask And Resize, which returns images, masks and bbox, if you intend to round-trip.
Trimming dead space before a VAE encode. A mask you drew on a large canvas often has a lot of irrelevant border; cropping to the box means the sampler spends latent budget where the content is.
Inspecting what a mask actually covers. Crop-only is a fast way to see whether your detector caught the face or the whole person.
Install
cd /path/to/ComfyUI/custom_nodes
git clone https://github.com/aining2022/ComfyUI_Swwan.git
cd ComfyUI_Swwan
python -m pip install -r requirements.txt
Windows portable:
.\python_embeded\python.exe -m pip install .\ComfyUI\custom_nodes\ComfyUI_Swwan\requirements.txt
Restart ComfyUI, hard-refresh the browser, search Swwan - this one is under Swwan/Advanced/Image. It ships in the pack's core image layer and needs no detection models: it never runs a detector, it just consumes whatever mask your detector produced. That's the pack's stated design across its whole masking layer, and it's why installing it doesn't drag in ultralytics or any other AGPL detector stack.
Requirements are just what requirements.txt pulls (opencv-python, scipy, scikit-image) on top of the torch, numpy and Pillow your ComfyUI already has. Python ≥3.11.
If a workflow comes up with a red missing node after a pack update, that's the 1.0.0 namespace migration, not a broken install:
python scripts/migrate_workflow.py old.json --dry-run
It writes to a new file by default, so the dry run costs you nothing but a second.
Inputs (2)
| Name | Type | Default | Description |
|---|---|---|---|
| image | IMAGE | — | |
| mask | MASK | — |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| image | IMAGE | — |