CV BBoxes to Masks
Turning boxes into masks (and knowing when not to)
- space
- bboxes
- masks
- count
Detection output is boxes, and masking wants masks. This is the one-step conversion: filled rectangles, one full-size mask per box, white inside. It's the inverse of CV Masks to BBoxes, and it's the piece that lets a box-emitting detector (YuNet, YOLO via RT-DETR, Match Template, SAM3's boxes) drive anything mask-shaped - inpainting, CV Crop by Masks, mask arithmetic, composite nodes.
The honest framing first: a rectangle is a blunt mask. If you're re-rendering a face, a box mask includes a chunk of hair and a chunk of wall, and those get redrawn too. The KB's detection doc has the clean version of this argument from someone who trains the segmentation face detectors - a segmentation model gives you a polygon, and a polygon creates a far less obvious seam. So use this when the rectangle is genuinely what you want: a quick region restriction, a keep-out zone, a coarse inpainting crop, or a starting mask that you'll grow or shrink downstream. If you need the seam to disappear, you want a segmentation detector, not a box.
Inputs
space- the reference the boxes were measured against. Only its height and width get used, so any IMAGE, MASK, NPARRAY or LATENT works. That's the whole reason it's a separate input: a box from a detector that ran on a downscaled copy has to be interpreted in the right coordinate space, and for a LATENT you're in the latent-cell grid, not pixels. CV Scale BBoxes is what converts between the two. Hand it the wrong reference and your masks land somewhere plausible but wrong - this is the number one mistake with this node.bboxes- coreBOUNDING_BOXdata, i.e. per-frame lists of{x, y, width, height}dicts. Anything that emits that type will do, including the core Bounding Box primitives, so it composes with detectors from other packs.merge- union every box into a single mask instead of emitting one mask per box. Off by default. Turn it on for "mask all the faces in this shot" jobs where you don't care which is which.
Outputs
masks is a MASK batch - one full-size mask per box, or one merged mask - and count tells you how many you got. Wire the batch into an inpainting node and the iteration is handled for you; each mask is its own job.
Empty input yields an empty MASK batch rather than raising. That's consistent across this pack's detectors and converters, and it's the right call for batch workflows: a frame with no detections shouldn't kill the run.
Install
Manager → ComfyUI CV, or:
cd ComfyUI/custom_nodes && git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"
Python ≥ 3.12, recent ComfyUI on the V3 node API. No models, no extra downloads. The pack's own tests run this node against core box emitters, so it stays honest when core changes the shape of BOUNDING_BOX - which, given how many places now emit it, is the actual risk with this type.
Where people get burned
Frame counts. If you have 8 boxes on one frame, you get 8 masks, and downstream nodes expecting one mask per image will happily process them in whatever order they arrive. When a mask count and an image count have to line up, be explicit about it - group_index from CV BBoxes To Array and CV Array To BBoxes exist for exactly that bookkeeping.
Soft edges. These masks are hard-edged by construction. Any inpainting workflow that expects a feathered or blurred mask will show the rectangle. The usual fix is a blur or dilation on the mask after this node, not a setting on it.
Inputs (3)
| Name | Type | Default | Description |
|---|---|---|---|
| space | NPARRAY,IMAGE,MASK,LATENT | Reference the boxes were measured on - only its height/width are used, so any IMAGE, MASK, NPARRAY or LATENT works. For a LATENT that is the latent-cell grid ('CV Scale BBoxes' converts between spaces). | |
| bboxes | BOUNDING_BOX | [object Object] | Core BOUNDING_BOX data: per-frame lists of {x, y, width, height} dicts - compatible with Draw BBoxes, Crop By Bounding Boxes, Image Crop, etc. |
| mergeopt | BOOLEAN | false | Union every box into ONE mask instead of emitting one mask per box. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| masks | MASK | One full-size mask per box (or a single merged mask), white inside the rectangle. |
| count | INT | — |