Nodes/ComfyUI CV/CV Crop by BBoxes
ComfyUI Node

CV Crop by BBoxes

40 regions, one node, and a lossless paste back

By bmad4ever·Created 4 months ago·Updated 15 days ago· 1
CV Crop by BBoxes
  • source
  • bboxes
  • crops
  • bboxes
  • width
  • height
  • count
◄policyone source, regions -> batch►

What it's for

The moment you have detection boxes, you want to do something per region: refine each face, inpaint each blur, upscale each crop. Core ComfyUI gives you Crop By Bounding Boxes and a prayer; the honest version of that workflow is a loop, and ComfyUI doesn't really have loops.

This node is the loop. Give it a BOUNDING_BOX and a source, get one crop per box in a single execution, run your region process once on the batch, paste back with CV Paste by BBox. No for-loop, no bookkeeping.

How it works

The important mechanism is that the crops are exact slices in the source's container type. IMAGE, MASK, NPARRAY and LATENT all work, and the crops come back as the type that went in. There's no uint8 round trip in the middle - no converting a latent to a picture to crop it, no clipping a float mask through 8-bit. That's what makes crop → process → paste genuinely lossless instead of lossy-with-a-straight-face.

The second mechanism is the policy dropdown, which decides which of four shapes you get:

  • one source, regions → batch - crops every region out of frame 0 at one common size and stacks them. This is the inpainting case: the downstream runs once on the batch.
  • one source, regions → list - one item per region, each at its own tight size, so the downstream runs once per region.
  • per-frame variants - pair region group g with frame g of a batch, instead of always reading frame 0.

The policy lives behind an advanced string input, so in the UI you get a mode dropdown plus a widget per option. Wire CV Snap BBoxes in and the widgets hide - one node then drives the geometry of several crops at once, and a linked policy only overrides the fields it actually sets.

The inputs and outputs that matter

  • source - the anything-typed input: IMAGE / MASK / NPARRAY / LATENT. Several inputs, or a batch, flatten into one frame sequence.
  • bboxes - core BOUNDING_BOX data: per-frame lists of {x, y, width, height}. It's compatible with Draw BBoxes, Crop By Bounding Boxes, Image Crop and friends, so it drops into existing graphs. The per-frame nesting is the grouping, which is why there's no separate "box list" input - wire several BOUNDING_BOX values in and they concatenate.

Out comes crops (the batch or the list), bboxes - the snapped boxes actually cropped, fanned out in lockstep with the crops - plus width, height and count as plain integers, so you can size a downstream node without guessing.

That second output is the one people miss. Feed CV Paste by BBox the boxes that came out of the crop, not the ones you put in. If a policy snapped or clamped anything, the input boxes are a lie about where the crop came from, and you get hairline seams.

Boxes are always clamped inside the source, and an empty input yields an empty batch rather than an exception.

Install

ComfyUI Manager, search ComfyUI CV. By hand:

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv

Restart after. The pack requires Python ≥ 3.12, a recent ComfyUI built on the V3 node API, and the contrib OpenCV wheel:

pip install "opencv-contrib-python-headless~=5.0.0.93"

Where people get burned

The latent coordinate trap. Boxes are in the source's coordinate space. Working on a pixel-space detection of a 1024×1024 image and then cropping the matching latent means your boxes are off by the VAE stride - 8×, in the usual case. CV Scale BBoxes exists for exactly that, and the node's own tooltip points you at it.

Batch mode pads, list mode doesn't. In the batch modes every crop is a common size (the largest one), which is what lets the paste line back up. If you want each region at its own tight size, you want list mode - and then you're running the downstream once per region, so be sure you meant to.

Lossless only if you stay in one type. The moment you convert a crop to a ComfyUI IMAGE and back, you've paid the quantisation. Cropping a latent, doing arithmetic on it as an NPARRAY and pasting it back is exact; routing it through a picture is not.

Nothing happens, quietly. Empty boxes in, empty batch out. If your downstream node reports zero images and nothing errored, look upstream at the detector, not here.

One pack-wide caveat, because it's the same for every node here. The boxes are core BOUNDING_BOX, so this node plays fine with core crop/paste/draw nodes - but the pack needs its contrib OpenCV intact. Installing a non-contrib wheel over it strips the contrib submodules and takes a chunk of the node list with them.

Categoryimage/CV

Inputs (3)

NameTypeDefaultDescription
sourceCOMFY_MATCHTYPE_V3What to crop. IMAGE / MASK / NPARRAY / LATENT; the crops come back in the same type. A batch - or several inputs - flattens into one frame sequence.
bboxesBOUNDING_BOX[object Object]Core BOUNDING_BOX data: per-frame lists of {x, y, width, height} dicts - compatible with Draw BBoxes, Crop By Bounding Boxes, Image Crop, etc. Boxes in the source's coordinate space. The per-frame nesting IS the grouping; several values are concatenated.
policyoptSTRINGone source, regions -> batchWhich use case this crop serves. 'one source, regions -> batch' crops every region out of frame 0 at one common size and stacks them (the inpainting case - downstream runs once on the batch). '-> list' instead emits one item per region at its own tight size, so downstream runs once per region. The 'per-frame' modes pair region group g with frame g of a batch instead of always reading frame 0. Crop policy. In the UI this is a mode dropdown plus one widget per option; wire an 'CV Snap BBoxes' node in to drive several crop nodes from one place (the widgets then hide). A linked policy overrides only the fields it actually sets. Boxes are always clamped inside the source.

Outputs (5)

NameTypeDescription
cropsCOMFY_MATCHTYPE_V3One stacked batch (batch modes) or one item per region (list modes), in the source's own type.
bboxesBOUNDING_BOXCore BOUNDING_BOX data: per-frame lists of {x, y, width, height} dicts - compatible with Draw BBoxes, Crop By Bounding Boxes, Image Crop, etc. The SNAPPED boxes actually cropped, fanned out in lockstep with 'crops' - feed these to 'OpenCV Paste by BBox', not the input boxes.
widthINTCrop width (the largest one when sizes vary).
heightINTCrop height (the largest one when sizes vary).
countINTNumber of crops (regions), whatever the output shape.