Nodes/ComfyUI CV/CV Saliency
ComfyUI Node

CV Saliency

A subject mask with no model file and no download

By bmad4ever·Created 4 months ago·Updated 14 days ago· 1
CV Saliency
  • image
  • saliency
  • binary
◄methodfine-grained►

Segmentation in ComfyUI usually means a model: SAM, rembg, a BiRefNet download, a checkpoint in the right folder. This node asks a different question - "what in this image is statistically unusual?" - using cv2.saliency, and answers it in milliseconds on CPU with nothing installed. The result is a rough mask, not a cutout. Knowing which of those you need is the whole decision.

The two methods

method switches between two classic detectors, and they fail in different directions:

  • fine-grained (the default) works on local centre-surround contrast at multiple scales. It follows object boundaries noticeably better, which is why it's the one to use when the output becomes a mask.
  • spectral-residual is the FFT log-spectrum trick - fast and global, highlighting whatever is unusual in the frequency domain. Coarser, and it tends to produce a general "interesting region" blob rather than a silhouette. Great as a cheap first pass or as an attention heatmap.

Neither is learned. Neither knows what a person is. On a high-contrast, centred subject they're surprisingly decent; on a busy scene with a bright background they'll happily pick the wrong thing at full confidence.

Outputs, and how to use them for real

saliency is a 0-1 MASK - how salient each pixel is - and it's genuinely useful as a soft weight rather than as a mask: wire it into a blend, multiply it against something, use it to steer where a crop should centre. CV Color Map turns it into a heatmap if you want to look at it.

binary is the auto-thresholded 0/1 version (Otsu-style, computed inside OpenCV). This is the mask-shaped output, and "empty" is a legitimate result on a flat, low-contrast image - the node returns an empty mask rather than raising, because a detector shouldn't halt a workflow.

Practical pattern: saliency for the rough region, then hand that to something with an actual opinion about edges. The pack's workflows/25_saliency_superpixels.json does exactly this - saliency plus the ximgproc superpixel family to build a subject mask without a model - and describes the two saliency modes in the same terms: fine-grained for masks, spectral-residual for speed. Also note the failure-tolerant design across this pack: if a detector returns nothing usable, you get zeros, not an exception.

Where it genuinely competes

Attention-weighted cropping is the honest sweet spot: find the salient region, crop around it, feed the crop to a detailer or an img2img pass. Same for auto-framing and for the "which part of this panorama should the thumbnail use" problem. It's also a solid interactive mask seed - paint a rough blob with saliency, then refine it by hand or with a second pass, instead of starting from a blank canvas.

Where it doesn't compete: extraction. If you need hair, semi-transparent edges, or a mask someone would ship, you want a learned matting model - the KB's own framing is that BiRefNet-class models produce significantly sharper edges, especially on hair and semi-transparent material. Saliency is the zero-cost first move, not the finish.

Install

Ships in ComfyUI CV (bmad4ever/comfyui_cv), a GPL-3.0 fork of opencv-comfyui:

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"
# restart ComfyUI

Manager users: search the pack title. Python ≥ 3.12 and a V3 node API ComfyUI. This node is exactly why the contrib wheel is not optional: cv2.saliency is a contrib module, and the pack's docs are explicit that installing a plain opencv-python over a contrib one silently empties the contrib submodules - the nodes just disappear with no error message. There's a repair script (tools/repair_opencv_contrib.py --check / --apply) but no install-time guard, so keep an eye on your own pip output.

Mechanism note

It's tagged CG in the pack docs: the entry point is a class (StaticSaliencyFineGrained_create), not a function, so the auto-generated cv2.* wrappers can't reach it - this curated node is the only route from a graph. It also catches the case where spectral-residual's binarization isn't implemented in your OpenCV build and falls back to Otsu on the saliency map itself. Small thing, and the kind of detail that saves you a confusing afternoon on a different wheel.

Categoryimage/CV/contrib

Inputs (2)

NameTypeDefaultDescription
imageNPARRAY,IMAGEImage to analyse (3-channel). Frame 0 of a batch. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
methodCOMBOfine-grained'fine-grained' follows object boundaries more closely; 'spectral-residual' is faster and coarser, highlighting whatever is unusual in the frequency spectrum.

Outputs (2)

NameTypeDescription
saliencyMASK0-1 MASK: how salient each pixel is. Feed it to 'CV Color Map' for a heatmap, or use it directly as a soft weight.
binaryMASKAuto-thresholded 0/1 MASK of the salient region (Otsu-style, computed by OpenCV). Empty is a valid result on a flat image.