CV DNN Mask Output
Turning a segmentation blob into an actual ComfyUI MASK
- blob
- image
- mask
Why you'd reach for this
A segmentation network run through cv2.dnn hands you a raw float blob: (1, 1, H, W), logits or probabilities or something in between, at whatever resolution the model works in - not your image's. ComfyUI wants a MASK: single-channel float32, in 0–1, at the size of the thing you're about to composite.
This node is the adapter. It squeezes the blob down to an H×W map, optionally squashes logits through a sigmoid, resizes to the image you actually fed the model, clips to 0–1 and produces a MASK. That's it, and that's the job: the difference between "look, numbers" and a mask you can plug into Overlay Masks, an inpainting graph, or anything else with a MASK input.
Be honest about the use case, though. Background removal is commoditised - ComfyUI shipped BiRefNet in core in May 2026, and there are several dedicated packs. You reach for this node when the model you want is only available as ONNX, or when you're already inside a cv2.dnn graph and don't want a second runtime in the same workflow.
How it works
The blob is squeezed to a 2-D map. Then sigmoid decides the squashing: auto applies a sigmoid only when the values fall outside roughly [0, 1] (which is what raw logits look like), yes always applies it, no never does. The 500-clip is applied before the exponential so a wild logit can't produce an overflow warning.
Target size comes from the image input if it's connected - a ComfyUI IMAGE tensor's H/W is read directly - otherwise from the width/height widgets, where 0 means "leave the model's native resolution alone". If a target exists and differs, it's a plain linear cv2.resize. Finally: clip to [0, 1], cast to float32, unsqueeze into a [1, H, W] tensor. One mask, not a batch.
Inputs and outputs that matter
- blob - the output of
CV DNN Forward. For a single-channel model that's(1, 1, H, W). Pair the loader withengine = classicfor Swin-Transformer backbones like BiRefNet; those are the exports the classic engine is there for. - sigmoid - leave it on
autoon your first run and check the result. If your mask comes out as a grey rectangle, the model was already emitting probabilities andautoguessed wrong; flip it tono. If it comes out looking like noise with a few white pixels, it's logits and you wantyes. - image - the original image you fed the model, so the mask lands on the right resolution. This is the one people forget, and the failure is silent: unconnected, the mask stays at the model's native size and any downstream composite is misaligned.
- mask - float32 MASK in
[0, 1]. Straight intoOverlay Masks, a mask-to-image bridge, or your inpainting graph.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"
Python ≥ 3.12, current ComfyUI (V3 node API), or install via ComfyUI Manager by searching ComfyUI CV. The .onnx model goes in ComfyUI/models/onnx; nothing is bundled, and model_sources.txt at the repo root lists where to get each model plus its license.
Common issues
- Mask is pure grey/fog. Wrong
sigmoidsetting for the model's output range. Tryyesandnoand see which one gives you an actual silhouette. - Mask is the right shape but shifted or scaled. You connected
imagefrom somewhere other than the tensor you actually sent to the model - a preprocessed or letterboxed copy. Seg models driven through a letterbox need the same padding undone before this node sees the output. - ONNX conversion is slower than the PyTorch original. It usually is. The KB's own note on BiRefNet is that the ONNX export runs something like 90% slower than the Swin-Large weights in
transformers- so convert for portability, not for speed. If you just want a cutout, use the core/CUDA path. - Contrib trap. Installing
opencv-pythonover the contrib wheel empties the contrib submodules and this pack's contrib-backed nodes disappear.tools/repair_opencv_contrib.py --checkdiagnoses it.
Inputs (5)
| Name | Type | Default | Description |
|---|---|---|---|
| blob | NPARRAY | Raw output from 'CV DNN Forward'. Typically (1, 1, H, W) for a single-channel segmentation model. | |
| sigmoid | COMBO | auto | 'auto' applies sigmoid only when values lie outside [0, 1] (i.e. raw logits). 'yes' always applies it. 'no' skips it (model already outputs probabilities). |
| imageopt | NPARRAY,IMAGE | Original input image, used to read the target resize dimensions. If disconnected, 'width' and 'height' widgets are used instead (0 = keep the model's native resolution). Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| widthopt | INT | 00–8192 | Target mask width in px. 0 = keep the model's output resolution (or use the connected image's width). |
| heightopt | INT | 00–8192 | Target mask height in px. 0 = keep the model's output resolution (or use the connected image's height). |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| mask | MASK | Float32 MASK in [0, 1] at the target resolution. Feed into 'Overlay Masks' or any mask consumer. |