Nodes/ComfyUI CV/CV YOLO Seg Masks
ComfyUI Node

CV YOLO Seg Masks

YOLO Instance Segmentation Without the Ultralytics Install

By bmad4ever·Created 4 months ago·Updated 14 days ago· 1
CV YOLO Seg Masks
  • det_rows
  • proto
  • image
  • masks
  • mask_count
◄ratio1.000►
◄pad_left0►
◄pad_top0►
◄net_size640►

You have a YOLO segmentation model, you have a photo, and what you actually want is one mask per object - the dog, the guy on the left, the bike rack - each cut to its outline. CV YOLO Seg Masks is the last link in that chain. It's the node where raw model output stops being a tensor and starts being something you can overlay, crop, or paste with.

Why you'd take this route at all

The usual ComfyUI answer for per-object masks is the Impact Pack: detector → SEGS → detailer. It works, but it runs on ultralytics weights, which means the ultralytics package in your environment plus an AGPL question that follows the model files around, not just the code.

This pack takes the other road: it runs the model itself, as ONNX, through cv2.dnn. No ultralytics install, no SEGS, no .pt file. What you pay instead is that you wire the pipeline yourself - letterbox, blob, forward pass, pick outputs, decode, and then this node. That's the deal.

What it actually does

A YOLO segmentation head emits two things. First, detection rows where the columns past the class id aren't detection metadata at all - they're mask coefficients, C of them per row. Second, a mask-prototype map of shape (1, C, mh, mw): a low-resolution basis, something like 160×160 for a 640 input.

Every instance mask is a linear combination of those prototypes, weighted by that row's coefficients. So the node computes logits = coeffs @ proto, reshapes back to mh × mw, zeroes everything outside that detection's box (in proto space, which is why one object's mask never leaks onto another's), bilinear-upsamples to the letterboxed network size, thresholds at 0 - the same as sigmoid > 0.5, since sigmoid(0) is 0.5 - trims the letterbox padding, and nearest-resizes the rest back to your original image. It's a faithful port of ultralytics' ops.process_mask, and the source reads that way line for line.

The output is a single (N, H, W) MASK batch at original resolution - a real ComfyUI MASK, so Overlay Masks, MaskToImage, or a detailer's mask input all take it directly.

The inputs that matter

  • det_rows - the (N, 6+C) rows from CV YOLO Detect Decode. Only the NMS-free layout ("rows: xyxy + conf + class id") carries coefficients through; a detection-only model or the wrong layout raises a clear error rather than producing empty masks.
  • proto - the prototype map, (1, C, mh, mw) or (C, mh, mw). You pick it out of the forward pass with CV DNN Pick Output; if its channel count doesn't match the coefficients per row, the node tells you so.
  • image - the original image, before letterboxing. Its size is the mask target size, so wire the same image you fed the letterbox node, not the padded one.
  • ratio, pad_left, pad_top, net_size - the numbers CV DNN Letterbox emitted, copied across. Defaults (1.0 / 0 / 0 / 640) are only correct if you never letterboxed. Feed a letterboxed image with default settings and the masks come out shifted and mis-scaled with no error at all - this is the number one way to make this node look broken when it isn't.
  • Outputs are masks and mask_count, which is just the detection count.

One limitation worth knowing before you build around it: it handles a single image, the first frame of a batch, same as the decode node it follows. Per-frame video looping is on you.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"

ComfyUI Manager: search the pack title comfyui_cv. You need Python ≥ 3.12 and a recent ComfyUI on the V3 node API. The contrib part of that wheel is not optional - installing plain opencv-python over it silently empties the contrib submodules and nodes vanish from the menu.

Models are not bundled. The pack's own seg workflow runs yolo26n-seg.onnx from the OpenCV contribution repo, dropped in ComfyUI/models/onnx, with a COCO label text file alongside it. Heads up: that model is AGPL-3.0 and the network clause is real - serving its output over a network pulls the obligations in. If that matters, bring a different export and check its head format against the decode layouts first.

Where people get burned

Beyond the letterbox numbers: don't expect masks out of a detection-only model - it can't, and says so. Zero detections is a normal result, not a failure; you get an empty mask batch. And take the pack's own warnings literally: heavy LLM assistance, possible overfitting elsewhere in the codebase, no planned updates, no support. Read it before you put it in anything that matters.

Categoryimage/CV/dnn

Inputs (7)

NameTypeDefaultDescription
det_rowsNPARRAYThe (N, 6+C) 'det_rows' from 'CV YOLO Detect Decode' (NMS-free seg layout): [x1, y1, x2, y2, conf, class_id, C mask coefficients] in LETTERBOXED pixels.
protoNPARRAYMask-prototype map, (1, C, mh, mw) or (C, mh, mw). C must equal the number of mask coefficients per det row.
imageNPARRAY,IMAGEThe ORIGINAL image (before letterboxing) - its size is the mask target size. An IMAGE batch uses its first frame. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
ratioFLOAT1.0000.000001–100Letterbox resize ratio from 'CV DNN Letterbox' (1.0 if the image was not letterboxed).
pad_leftINT00–4096Letterbox left padding from 'CV DNN Letterbox'.
pad_topINT00–4096Letterbox top padding from 'CV DNN Letterbox'.
net_sizeINT64032–4096Square letterboxed input size the model ran at (the 'size' used in 'CV DNN Letterbox').

Outputs (2)

NameTypeDescription
masksMASK(N, H, W) binary instance masks at the original image size, aligned row-for-row with the Detect Decode bboxes. Feed 'Overlay Masks' or any mask consumer.
mask_countINTNumber of masks built (= the detection count).