CV YOLO Detect Decode
Turning a raw YOLO head into boxes nobody has to guess about
- det
- image
- confs
- bboxes
- scores
- class_ids
- labels_out
- det_rows
- det_count
- found
A YOLO ONNX export doesn't output boxes. It outputs a tensor - rows of numbers in a format that varies by export script, YOLO version, and the mood of whoever wrote the converter. Turning that tensor into "three people, at these pixel coordinates, with these confidences" is a decoding problem, and getting it wrong is how you end up with boxes that are all in the top-left corner or scaled to the wrong resolution.
CV YOLO Detect Decode is that decoder. Point it at the head output, tell it which layout you've got, and you get boxes in original-image pixels plus scores, class ids and labels.
The four layouts
The layout dropdown covers the head formats that actually turn up:
- rows: xyxy + conf + class id (NMS-free) - the default, and what yolo26/yolov11-style end-to-end heads emit: rows of
[x1, y1, x2, y2, conf, class_id, ...]in letterboxed pixels. Nothing to suppress, because the head already did it. Any extra columns ride along untouched. - rows: xywh center + class scores (NMS) - yolov5/v8 heads with per-class score columns and centre-form boxes. NMS is applied.
- columns: xywh center + class scores (NMS) - the same data transposed, which is a genuinely different tensor shape depending on the exporter.
- separate boxes + confs (NMS) - yolov4-style exports: normalised xyxy boxes in one output and per-class confidences in another. That second tensor is the confs input, and it's required for this layout; it's ignored otherwise.
The description also makes the segmentation point cleanly: this node does no mask handling at all. That's deliberate - it means a detection-only model works here too, and if you do want instance segmentation you chain CV YOLO Seg Masks afterwards. The det_rows output is what makes that possible: (N, 6+K) float32 in letterboxed pixels, [x1, y1, x2, y2, conf, class_id, ...], with the NMS-free seg layout's mask coefficients riding along in the extra columns.
Inputs that matter
- det - the head output. Its expected shape depends entirely on
layout, which is why that dropdown is the first thing to get right. - image - the original image, before letterboxing. This is the coordinate target: the node un-letterboxes boxes back to original-image pixels, so it needs to know the original size.
- ratio, pad_left, pad_top - the letterbox parameters, straight from
CV DNN Letterbox. Leave them at defaults only if you didn't letterbox.net_size(default 640) is used solely by the separate-boxes layout, whose boxes are normalised to[0, 1]of it. - conf_threshold (default 0.5), nms_threshold (default 0.45, NMS layouts only), and labels - class names one per line, line N = class id N, e.g. piped from
CV Load Labels. Blank and detections are labelled by their numeric id.
Outputs: bboxes (a BOUNDING_BOX per detection with score and label - straight into the core Draw BBoxes), scores, class_ids, labels_out (strings, for Preview as Text), det_rows, det_count, and found. DATA only, you draw it yourself, and zero detections is a valid result. Gate the visualisation branch on found.
The licensing thing, since this is YOLO
The pack's model-sourcing file is blunt: the example segmentation model, yolo26n-seg.onnx, is AGPL-3.0 - copyleft, with a network clause. Apache-2.0 covers the packaging only; Ultralytics' terms govern the weights. Nothing here redistributes it, but serving its output over a network pulls the obligations in. The YOLO path also has a live supply-chain history - a poisoned Ultralytics release shipped a cryptominer that reached ComfyUI users through the detailing node pack in late 2024 - so if your use is commercial, the KB's detection essay names the licence-clean alternative (a MediaPipe detector, Apache-2.0).
Installing it
Manager → ComfyUI CV, or:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
Restart. Python ≥ 3.12, current ComfyUI on the V3 node API, opencv-contrib-python-headless~=5.0.0.93 - the whole pack runs cv2.dnn, so the wheel version is your inference runtime. Models go in ComfyUI/models/onnx and don't ship with the repo. GPL-3.0, forked from opencv-comfyui.
Where people get burned
Wrong layout. Boxes come out as nonsense - everything in one corner, coordinates in the thousands, or a handful of absurdly large boxes. The layout is a guess about a tensor format, so the fastest fix is to print the head output's shape (via Inspect CV Data) and match it against the shapes listed in the det tooltip.
Letterbox parameters left at defaults. You padded and resized to 640, and the default ratio = 1, pad_left = 0 means the boxes come back offset and scaled. Wire all three from the same CV DNN Letterbox node that produced the letterboxed input. This is the single most common "the boxes are almost right" complaint.
A batch. Only the first frame is used. For video, run per frame.
Expecting masks. They're a separate node on purpose - CV YOLO Seg Masks, fed from det_rows.
Comparing against the native YOLO node. ComfyUI's ecosystem runs YOLO in PyTorch on the GPU, from .pt files, with no conversion. This path goes through cv2.dnn instead - ONNX export, correct layout, CPU inference - which is the pack's whole premise. Use it to see the pieces fit; use the native node when you want speed.
Inputs (11)
| Name | Type | Default | Description |
|---|---|---|---|
| det | NPARRAY | Detection head output. Shape depends on 'layout': (1, N, 6+C)/(N, 6+C) rows for the NMS-free layout, (N, 4+C) or (1, 4+C, N) for the class-scores layouts, (1, N, 1, 4) normalized xyxy for the separate-boxes layout. | |
| image | NPARRAY,IMAGE | The ORIGINAL image (before letterboxing) - its size is the bbox target size. An IMAGE batch uses its first frame. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| ratio | FLOAT | 1.0000.000001–100 | Letterbox resize ratio from 'CV DNN Letterbox' (1.0 if the image was not letterboxed). |
| pad_left | INT | 00–4096 | Letterbox left padding from 'CV DNN Letterbox'. |
| pad_top | INT | 00–4096 | Letterbox top padding from 'CV DNN Letterbox'. |
| net_size | INT | 64032–4096 | Square letterboxed input size the model ran at. Only used by the 'separate boxes + confs' layout, whose boxes are normalized to [0, 1] of it. |
| conf_threshold | FLOAT | 0.500–1 | Minimum confidence to keep a detection. Lower detects more (and more false positives). |
| layout | COMBO | rows: xyxy + conf + class id (NMS-free) | The detection head's row/column format (see the node description). |
| confsopt | NPARRAY | Per-class confidence tensor (1, N, C) - REQUIRED for the 'separate boxes + confs' layout, ignored otherwise. | |
| nms_thresholdopt | FLOAT | 0.450–1 | Non-maximum suppression IoU, used by the NMS layouts: boxes overlapping by more than this are merged. |
| labelsopt | STRING | Class names, one per line (line N = class id N), e.g. from 'CV Load Labels'. Blank = detections are labelled by their numeric class id. |
Outputs (7)
| Name | Type | Description |
|---|---|---|
| bboxes | BOUNDING_BOX | One {x, y, width, height, score, label} dict per detection in ORIGINAL-image pixels - feed the core 'Draw BBoxes' node. |
| scores | NPARRAY | (N,) float32 confidence per detection, same order as bboxes. |
| class_ids | NPARRAY | (N,) int32 class id per detection, same order as bboxes. |
| labels_out | NPARRAY | (N,) string array of the class label per detection (the name from 'labels', or the numeric id). Feed 'Preview as Text'. |
| det_rows | NPARRAY | (N, 6+K) float32 kept rows in LETTERBOXED pixels [x1, y1, x2, y2, conf, class_id, ...extra columns from the head - the NMS-free seg layout's mask coefficients ride along here]. Feed 'CV YOLO Seg Masks'. Empty (0, 6) when nothing is detected. |
| det_count | INT | Number of detections kept. |
| found | BOOLEAN | True when at least one detection passed the confidence threshold. Feed an 'if/else' node to skip downstream visualization when empty. |