CV BBoxes To Array
Reading BOUNDING_BOX data as plain arrays
- bboxes
- boxes
- scores
- label_ids
- labels
- group_index
- count
Here's a wall people hit the first time they try to do anything clever with detections. Every detector in ComfyUI emits a BOUNDING_BOX - a list of {x, y, width, height} dicts, sometimes with score and label keys - and that type is deliberately opaque. You can hand it to Draw BBoxes, to Crop By Bounding Boxes, and not much else. You cannot sort it. You cannot threshold it. You cannot compute on it. This node is the bridge: BOUNDING_BOX in, ordinary NPARRAY out, and now the entire numeric half of this pack is open to you.
It's the numeric bridge out of the detection world. CV Array To BBoxes is the return leg, so a round trip detects → measures → filters → reconstructs without ever leaving graphs.
It isn't picky about the source
The input is core BOUNDING_BOX data, which means it accepts output from any emitter, not just this pack: core RT-DETR, SAM3 nodes, MediaPipe nodes, Bounding Box, or boxes you authored by hand. A single dict, a flat list, or per-frame nested lists all parse. If the boxes came from several frames, the nesting is preserved in group_index, which is how you get back to the frame a row came from.
Inputs
bboxes- the BOUNDING_BOX value.layout- the column layout of the array side, which matters because there are three conventions in common use and they're not interchangeable:x, y, width, height(OpenCV's own, and what BOUNDING_BOX is),x1, y1, x2, y2(most DNN heads and torchvision), andcenter x, center y, width, height(YOLO). Only the array changes - BOUNDING_BOX is always x/y/width/height. Get this wrong and nothing errors; your boxes are just in the wrong place, which is much worse.
Outputs
boxes is (N, 4) float32 in your chosen layout. scores is (N,) float32 pulled from each box's score key - and boxes with no score read 1.0, so a downstream threshold never silently deletes un-scored boxes. label_ids is (N,) int32 indexing into labels, the string list of distinct label names, one per line (the same shape Load Labels produces); unlabelled boxes all share index 0. That pair is what you feed to CV NMS Boxes for per-class suppression.
group_index is the per-frame provenance, and count is the total row count.
Two conventions in this pack are worth knowing because they're not universal: scores are float32 (not float64), so a round trip keeps about 7 digits rather than all 17, and the node never raises - empty input gives empty arrays.
Typical uses
Thresholding and sorting detections that a detector gives you in an arbitrary order. Computing an overlap in frames where boxes move between two sources. Measuring box areas to drop the noise hits before a detailer pass. And feeding the whole thing into the raw cv2.* wrappers or the pack's plotting nodes, which take NPARRAY and nothing else.
The KB's take on the detector layer applies directly: the interesting problems aren't in the detection, they're in the filtering - which faces to keep, which duplicates to merge, which are too small to bother re-rendering. That work needs numbers, and numbers are what this node produces.
Install
Manager → ComfyUI CV, or:
cd ComfyUI/custom_nodes && git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"
Python ≥ 3.12 and a recent ComfyUI on the V3 node API. Nothing else to configure, and no model downloads.
Where people get burned
Precision mismatches: cast the result with CV Cast Array if a consumer wants int32, and remember that layout is a label on the columns, not a conversion. Feeding labels where a numeric index is expected (or vice versa) is the other common slip - label_ids is the numbers, labels is the names.
Inputs (2)
| Name | Type | Default | Description |
|---|---|---|---|
| bboxes | BOUNDING_BOX | [object Object] | Core BOUNDING_BOX data: per-frame lists of {x, y, width, height} dicts - compatible with Draw BBoxes, Crop By Bounding Boxes, Image Crop, etc. A single dict, a flat list or per-frame nested lists are all accepted; the frame each box came from is reported in 'group_index'. |
| layoutopt | COMBO | x, y, width, height | Column layout of the array. 'x, y, width, height' is OpenCV's own convention (and BOUNDING_BOX's); 'x1, y1, x2, y2' is the corner form most DNN heads and torchvision use; 'center x, center y, width, height' is the YOLO form. Only the ARRAY side changes - BOUNDING_BOX is always {x, y, width, height}. |
Outputs (6)
| Name | Type | Description |
|---|---|---|
| boxes | NPARRAY | (N,4) float32, one row per box in the chosen layout, in the order the boxes appear (frame by frame). Cast it with 'CV Cast Array' if a consumer wants int32. |
| scores | NPARRAY | (N,) float32 confidence, from each box's 'score' key. Boxes with no score read as 1.0, so a threshold never silently drops un-scored boxes. float32 like every other score in this pack, so a round trip through 'CV Array To BBoxes' keeps ~7 digits, not all 17. |
| label_ids | NPARRAY | (N,) int32 index into 'labels'. Boxes with no label all share index 0. Feed it to 'CV NMS Boxes' for per-class suppression. |
| labels | STRING | The distinct label names, one per line, indexed by 'label_ids' - the same shape 'Load Labels' produces. |
| group_index | NPARRAY | (N,) int32: which per-frame GROUP each row came from. Pass it back to 'CV Array To BBoxes' to restore the original nesting. |
| count | INT | Total number of boxes (rows). |