Nodes/ComfyUI CV/CV DNN Mask Output
ComfyUI Node

CV DNN Mask Output

Turning a segmentation blob into an actual ComfyUI MASK

By bmad4ever·Created 3 months ago·Updated 14 days ago· 1
CV DNN Mask Output
  • blob
  • image
  • mask
◄sigmoidauto►
◄width0►
◄height0►

Why you'd reach for this

A segmentation network run through cv2.dnn hands you a raw float blob: (1, 1, H, W), logits or probabilities or something in between, at whatever resolution the model works in - not your image's. ComfyUI wants a MASK: single-channel float32, in 0–1, at the size of the thing you're about to composite.

This node is the adapter. It squeezes the blob down to an H×W map, optionally squashes logits through a sigmoid, resizes to the image you actually fed the model, clips to 0–1 and produces a MASK. That's it, and that's the job: the difference between "look, numbers" and a mask you can plug into Overlay Masks, an inpainting graph, or anything else with a MASK input.

Be honest about the use case, though. Background removal is commoditised - ComfyUI shipped BiRefNet in core in May 2026, and there are several dedicated packs. You reach for this node when the model you want is only available as ONNX, or when you're already inside a cv2.dnn graph and don't want a second runtime in the same workflow.

How it works

The blob is squeezed to a 2-D map. Then sigmoid decides the squashing: auto applies a sigmoid only when the values fall outside roughly [0, 1] (which is what raw logits look like), yes always applies it, no never does. The 500-clip is applied before the exponential so a wild logit can't produce an overflow warning.

Target size comes from the image input if it's connected - a ComfyUI IMAGE tensor's H/W is read directly - otherwise from the width/height widgets, where 0 means "leave the model's native resolution alone". If a target exists and differs, it's a plain linear cv2.resize. Finally: clip to [0, 1], cast to float32, unsqueeze into a [1, H, W] tensor. One mask, not a batch.

Inputs and outputs that matter

  • blob - the output of CV DNN Forward. For a single-channel model that's (1, 1, H, W). Pair the loader with engine = classic for Swin-Transformer backbones like BiRefNet; those are the exports the classic engine is there for.
  • sigmoid - leave it on auto on your first run and check the result. If your mask comes out as a grey rectangle, the model was already emitting probabilities and auto guessed wrong; flip it to no. If it comes out looking like noise with a few white pixels, it's logits and you want yes.
  • image - the original image you fed the model, so the mask lands on the right resolution. This is the one people forget, and the failure is silent: unconnected, the mask stays at the model's native size and any downstream composite is misaligned.
  • mask - float32 MASK in [0, 1]. Straight into Overlay Masks, a mask-to-image bridge, or your inpainting graph.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"

Python ≥ 3.12, current ComfyUI (V3 node API), or install via ComfyUI Manager by searching ComfyUI CV. The .onnx model goes in ComfyUI/models/onnx; nothing is bundled, and model_sources.txt at the repo root lists where to get each model plus its license.

Common issues

  • Mask is pure grey/fog. Wrong sigmoid setting for the model's output range. Try yes and no and see which one gives you an actual silhouette.
  • Mask is the right shape but shifted or scaled. You connected image from somewhere other than the tensor you actually sent to the model - a preprocessed or letterboxed copy. Seg models driven through a letterbox need the same padding undone before this node sees the output.
  • ONNX conversion is slower than the PyTorch original. It usually is. The KB's own note on BiRefNet is that the ONNX export runs something like 90% slower than the Swin-Large weights in transformers - so convert for portability, not for speed. If you just want a cutout, use the core/CUDA path.
  • Contrib trap. Installing opencv-python over the contrib wheel empties the contrib submodules and this pack's contrib-backed nodes disappear. tools/repair_opencv_contrib.py --check diagnoses it.
Categoryimage/CV/dnn

Inputs (5)

NameTypeDefaultDescription
blobNPARRAYRaw output from 'CV DNN Forward'. Typically (1, 1, H, W) for a single-channel segmentation model.
sigmoidCOMBOauto'auto' applies sigmoid only when values lie outside [0, 1] (i.e. raw logits). 'yes' always applies it. 'no' skips it (model already outputs probabilities).
imageoptNPARRAY,IMAGEOriginal input image, used to read the target resize dimensions. If disconnected, 'width' and 'height' widgets are used instead (0 = keep the model's native resolution). Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
widthoptINT00–8192Target mask width in px. 0 = keep the model's output resolution (or use the connected image's width).
heightoptINT00–8192Target mask height in px. 0 = keep the model's output resolution (or use the connected image's height).

Outputs (1)

NameTypeDescription
maskMASKFloat32 MASK in [0, 1] at the target resolution. Feed into 'Overlay Masks' or any mask consumer.