CV Crop by Masks
Mask in, context-padded crops out
- image
- masks
- crops
- crop_masks
- bboxes
- region_index
What it's for
Masks are how ComfyUI talks about "this part of the image". Crops are how you actually process it, because diffusion on a 512×512 region is cheaper and sharper than diffusion on a 4K frame. The bridge between them has traditionally been masks → masks-to-bboxes → crop → remember which box went with which mask. Three nodes and a bookkeeping problem.
CV Crop by Masks is that whole bridge in one node. Give it a source and a MASK batch, get crops - plus the matching cropped masks, the bounding boxes, and an index that tells you which input mask each crop came from.
It's the mask-driven shorthand for CV Masks to BBoxes → CV Crop by BBoxes, and it takes the same policy - so one CV Snap BBoxes can drive this and a box-based crop in lockstep, instead of two crops disagreeing about geometry.
How it works
Each mask becomes a region. The region is grown by padding - 16 pixels by default - because inpainting works far better when the model can see the surroundings that the mask was cut out of. Set padding to 0 and you get a crop that ends exactly at the mask's own extent, which is the single most common reason a per-region inpaint comes back with a visible edge band.
The same policy vocabulary as CV Crop by BBoxes applies: batch modes stack every crop at one common size (so they paste back losslessly and the downstream runs once), list modes emit one item per region at its own tight size (downstream runs once per region), and the per-frame modes pair region group g with frame g instead of reading frame 0.
The source can be an IMAGE, a MASK, an NPARRAY or a LATENT, and crops come back as whatever went in. Masks are fitted to the source's size internally, which is why a LATENT source works fine with ordinary pixel-space masks - no manual scaling dance.
The inputs and outputs that matter
- image - slightly misleadingly named, but it's the anything-input: what you crop.
- masks - a
MASKbatch, one mask per region. Several batches wired in are concatenated. - padding (default 16, 0–4096) - the context ring around each region.
- square (default off) - forces square crops, which is the setting to reach for when the crops are going into a diffusion inpainting pass that expects square latents.
Four outputs. crops and crop_masks come out in lockstep - the cropped mask for each crop is the thing you'll feed an inpaint node, and getting it from the same node is why this exists. bboxes gives you the regions as core BOUNDING_BOX data, ready for paste or draw. And region_index is a (B,) int32 array telling you which row of the input mask batch each crop came from.
That last one sounds like trivia and isn't. Masks that are empty produce no crop at all. So if three of your twenty masks are blank, you get seventeen crops, and region_index is the only thing that tells you which seventeen.
Install
ComfyUI Manager, search ComfyUI CV, or:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
Restart, then:
pip install "opencv-contrib-python-headless~=5.0.0.93"
Python ≥ 3.12 and a recent ComfyUI (V3 node API) are the hard requirements.
Where people get burned
Padding is not optional in practice. The default of 16 exists because tight crops inpaint badly. For a face, 32–64 is a reasonable start; the wider the ring, the more the model knows about the lighting it's blending into.
Empty masks vanish. By design, and it means you can't blindly zip the crops back against your input list. Use region_index.
Paste with the node's own boxes. The bboxes output is what matches the crops. If a policy snapped anything and you paste using the input geometry instead, you get seams that look like a model problem and aren't.
Square crops change the geometry. square is right for diffusion inpaint pipelines and wrong if you plan to paste back with the same boxes - the crop is no longer the region the box describes.
Latent sources still take pixel masks. That's a feature, not a gotcha, but it's the source of the reverse confusion too: boxes you export for use elsewhere are in the source's own coordinate space, so a latent needs CV Scale BBoxes before anything pixel-based consumes them.
One pack-wide warning, and it's the author's own. This is a GPL-3.0 fork of geroldmeisinger's opencv-comfyui, written by one person with heavy LLM assistance, and the README says plainly it isn't production-ready without independent review and your own tests. Nodes are generated from whatever OpenCV build is installed, so a non-contrib opencv-python wheel landing on top of the contrib one empties the contrib submodules and takes part of this pack with it.
Inputs (5)
| Name | Type | Default | Description |
|---|---|---|---|
| image | COMFY_MATCHTYPE_V3 | What to crop. IMAGE / MASK / NPARRAY / LATENT; the crops come back in the same type. | |
| masks | MASK | One mask per region to crop (MASK batch). Several batches wired in are concatenated. | |
| padding | INT | 160–4096 | Extra context around each region in pixels - inpainting works better with surroundings included. |
| square | BOOLEAN | false | Force square crops (useful for diffusion inpainting). |
| policyopt | STRING | one source, regions -> batch | Which use case this crop serves. 'one source, regions -> batch' crops every region out of frame 0 at one common size and stacks them (the inpainting case - downstream runs once on the batch). '-> list' instead emits one item per region at its own tight size, so downstream runs once per region. The 'per-frame' modes pair region group g with frame g of a batch instead of always reading frame 0. Crop policy. In the UI this is a mode dropdown plus one widget per option; wire an 'CV Snap BBoxes' node in to drive several crop nodes from one place (the widgets then hide). A linked policy overrides only the fields it actually sets. The padding/square widgets above are the baseline; a policy overrides only the fields it sets. |
Outputs (4)
| Name | Type | Description |
|---|---|---|
| crops | COMFY_MATCHTYPE_V3 | One stacked batch (batch modes) or one item per region (list modes), in the source's own type. |
| crop_masks | MASK | Each region's mask cropped to its crop, fanned out in lockstep with 'crops'. |
| bboxes | BOUNDING_BOX | Core BOUNDING_BOX data: per-frame lists of {x, y, width, height} dicts - compatible with Draw BBoxes, Crop By Bounding Boxes, Image Crop, etc. Fanned out in lockstep with 'crops'. |
| region_index | NPARRAY | (B,) int32: which row of the input MASK batch each crop came from. Empty masks produce no crop, so this is what realigns the crops with the input batch. |