CV Array To BBoxes
How a numpy table becomes something ComfyUI can crop
- boxes
- scores
- label_ids
- group_index
- bboxes
- count
Detections arrive as numbers. ComfyUI's crop, draw and region nodes want BOUNDING_BOX - the structured type you get from a core PrimitiveBoundingBox or a detector node, with per-frame lists of {x, y, width, height} dicts. Something has to translate, and it's this node. It's the return leg of CV BBoxes To Array, and it's the reason any array of regions - a decoded DNN head, k-means cluster extents, boxes computed with Math Expression - can be cropped, drawn and pasted.
Why you'd reach for it
Because at some point you will build boxes yourself. A segmentation node that gives you cluster extents, a threshold pass whose connected components you measured, a template match, a hand-written numpy expression in a Math Expression node that computes a region from landmarks - all of those end up as rows in an array, and none of them is a BOUNDING_BOX.
Wire this in and the whole core surface opens up: Draw BBoxes, Crop By Bounding Boxes, Image Crop, plus everything in the pack that consumes boxes. Scores and labels ride along as extra keys, and every bbox node in the pack preserves them - so a threshold or a class name you attached here survives to whatever draws it at the end.
How it works, and why layout is the whole node
boxes is an (N,4) or (N,≥4) array, and the columns have to mean something - which is what layout tells it:
x, y, width, height- the default, and OpenCV's own convention (which is alsoBOUNDING_BOX's).x1, y1, x2, y2- the corner form that most DNN heads and torchvision emit.center x, center y, width, height- the YOLO form.
Only the array side changes; BOUNDING_BOX is always {x, y, width, height}. Get the layout wrong and you don't get an error, you get plausible-looking boxes in the wrong place - a x2, y2 corner read as a width and height produces boxes that grow with the frame instead of tracking it, which is the kind of thing people stare at for ten minutes. Extra columns beyond the fourth are ignored, so a detection table with a trailing attribute is fine, and a single (4,) row is accepted as one box. Floats are rounded, because boxes are pixels.
Optional inputs fill in the metadata:
scores- an(N,)array of confidences, written to each box'sscorekey.label_ids- an(N,)array of int indices intolabels; without alabelsstring, the index itself becomes the label.labels- one label name per line, which takes the output of a coreLoad Labelsnode directly.group_index- an(N,)int frame index per row. This is what restores per-frame nesting: give it the frame column you got fromCV BBoxes To Arrayand boxes go back into their frames; omit it and every box lands in one group.
Outputs are bboxes (BOUNDING_BOX) and count (INT). Empty input yields an empty BOUNDING_BOX and never raises, so a graph runs fine before any detections exist - which is the difference between a workflow you can open and a workflow that errors out on load.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
then restart ComfyUI, or use ComfyUI Manager → ComfyUI CV. opencv-contrib-python-headless~=5.0.0.93 is the declared dependency; this node itself needs only numpy. Python 3.12+ and a recent V3-API ComfyUI are required.
Traps
- Layout mismatches are silent. Check one box's numbers against the image before you trust a batch of them - a box wider than the frame or at a negative coordinate is the tell.
- Rounding is to integers. Sub-pixel boxes are a thing in detection; they're not a thing in
BOUNDING_BOX. Accumulate the error over a scaling step and you'll see it. - Omit
group_indexand you get one group. For a batch of frames, "one group" means the boxes are no longer associated with frames, and a per-frame crop node will do something you didn't intend. Feed the frame index through if you came from the array side. label_idsneed a matchinglabelsstring. Mismatched lengths are the classic source of a label that's just a number.- It's a type bridge, not a filter. Nothing here validates boxes against the image, deduplicates them or clamps them to the frame. Non-max suppression is a separate node (
CV NMS Boxes) for exactly that reason.
Inputs (6)
| Name | Type | Default | Description |
|---|---|---|---|
| boxes | NPARRAY | (N,4) or (N,>=4) array of box rows in the chosen layout; extra columns are ignored. A single (4,) row is accepted as one box. Floats are rounded. | |
| layoutopt | COMBO | x, y, width, height | Column layout of the array. 'x, y, width, height' is OpenCV's own convention (and BOUNDING_BOX's); 'x1, y1, x2, y2' is the corner form most DNN heads and torchvision use; 'center x, center y, width, height' is the YOLO form. Only the ARRAY side changes - BOUNDING_BOX is always {x, y, width, height}. |
| scoresopt | NPARRAY | Optional (N,) confidences, written to each box's 'score' key. Omit to leave the boxes un-scored. | |
| label_idsopt | NPARRAY | Optional (N,) int indices into 'labels'. Without 'labels' the index itself is written as the label. | |
| labelsopt | STRING | Optional label names, one per line, indexed by 'label_ids' - takes 'Load Labels' output directly. | |
| group_indexopt | NPARRAY | Optional (N,) int frame index per row (as emitted by 'CV BBoxes To Array'), restoring the per-frame nesting. Omitted, every box lands in one group. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| bboxes | BOUNDING_BOX | Core BOUNDING_BOX data: per-frame lists of {x, y, width, height} dicts - compatible with Draw BBoxes, Crop By Bounding Boxes, Image Crop, etc. |
| count | INT | Number of boxes emitted. |