Nodes/ComfyUI CV/CV HFS Segmentation (hierarchical feature selection)
ComfyUI Node

CV HFS Segmentation (hierarchical feature selection)

Superpixel regions with no model and no GPU

By bmad4ever·Created 3 months ago·Updated 14 days ago· 1
CV HFS Segmentation (hierarchical feature selection)
  • image
  • segmented
  • labels
  • num_regions
◄egb_threshold_i0.08►
◄min_region_size_i100►
◄egb_threshold_ii0.28►
◄min_region_size_ii200►
◄spatial_weight0.60►
◄superpixel_size8►
◄slic_iterations5►

Not every segmentation job is "find the person." Sometimes you want the structure of the image: which pixels belong together, where the boundaries between visually distinct regions are, how many regions there even are. That's what hierarchical feature selection does - SLIC superpixels merged by texture (texton) and colour similarity - and it runs on your CPU with no weights.

It's cv2.hfs wrapped as a node in bmad4ever's ComfyUI CV (bmad4ever/comfyui_cv, a fork of Gerold Meisinger's opencv-comfyui). It sits in the contrib corner of the pack for a reason: hfs lives in OpenCV's ximgproc module, so it's contrib-only and it's one of the many class-based APIs (HfsSegment_create) that the auto-generated raw cv2.* wrappers can't reach.

How it works

Two stages, and they're the reason the node has four thresholds instead of two. SLIC clusters pixels into compact superpixels - that's superpixel_size (8 px, smaller = more seeds = finer detail) and slic_iterations (5 is plenty; 10 on very textured content), with spatial_weight (0.6) trading proximity against colour/texture similarity inside the clustering. Lower it and colour dominates.

Then the pre-merged superpixels are merged hierarchically: gradient-based region merging under a graph-based criterion, twice. egb_threshold_i (0.08) and min_region_size_i (100) run the first, finer pass; egb_threshold_ii (0.28) and min_region_size_ii (200) run the second, coarser one - which is why the second threshold is comfortably larger. Higher thresholds merge more aggressively: fewer, larger regions, less detail retained. Region size floors absorb the stragglers.

The practical trade, stated plainly: region purity against region count. You're not optimising for "correct", you're optimising for the granularity you want downstream.

Outputs

  • segmented - the input image with each region repainted in its average colour. This is what you'll actually enjoy: a clean, poster-like abstraction where the structure of the image is legible. It's a legitimately nice stylization (deterministic, millisecond, no prompt), in the same family as the palette-quantisation moves the KB describes for pixel art.
  • labels - a uint16 per-pixel region index, 0 to num_regions - 1, same height and width as the input. That uint16 is the important detail: it's an index map, not a displayable image, so preview it with normalize/heatmap mode and expect a false-colour mess if you look at it raw.
  • num_regions - how many distinct regions came out. Your main tuning feedback signal; print it while you adjust thresholds.

labels is the output that goes to work. Feed it into CV Labels to Masks (full size) to get one MASK per region, then into masking, cropping or plotting nodes. That's the classical route to "give me masks for all the distinct areas" without a promptable segmenter - perfect for region-based colour work, measuring areas, or building a mask you'll refine by hand.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"

Or find ComfyUI CV in ComfyUI Manager. Python ≥ 3.12, V3-API ComfyUI. Because hfs is contrib, this node is a live test of whether your OpenCV install is clean: if the pack loads but this node is missing, a non-contrib wheel got installed over the contrib one and emptied the submodules. tools/repair_opencv_contrib.py --check will tell you. workflows/26_hfs_segmentation.json is a six-node workflow - load the coins photo, segment, preview - and it's the fastest way to get a feel for the thresholds.

Common issues

Region map looks like noise. superpixel_size too small combined with high iteration counts, or the merges never happened - check min_region_size_i/ii aren't so large they're swallowing everything into one region, and that the thresholds aren't at 0.

One giant region. Both thresholds too high. Dial egb_threshold_i down first; the first stage is where most of the detail is preserved or lost.

Everything is preserved, hundreds of regions. Thresholds too low. Raise egb_threshold_ii and increase the minimum sizes to collapse small regions into neighbours.

Slow on a 4K image. It's CPU and it scales with pixel count; slic_iterations and superpixel_size are the levers. Segment at working resolution and scale the label map, or accept the wait.

Don't expect semantic labels. A region is "these pixels look alike", not "this is a coin". For object-level masks you want the model-backed route - BiRefNet, SAM, or a promptable segmenter. This is the structural, model-free end of the same spectrum.

Categoryimage/CV/contrib

Inputs (8)

NameTypeDefaultDescription
imageCOMFY_MATCHTYPE_V3Full-colour image to segment (3 channels). The region index map has one label per pixel and the same height/width. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
egb_threshold_ioptFLOAT0.080–1Stage-I merge threshold: gradient below this and two regions merge. Higher = fewer, larger regions (and less detail kept).
min_region_size_ioptINT1001–100000Stage-I minimum region size in pixels; regions smaller than this get absorbed into a neighbour.
egb_threshold_iioptFLOAT0.280–1Stage-II threshold, applied to the coarser regions. Usually comfortably larger than egb_threshold_i, hence the second coarser pass.
min_region_size_iioptINT2001–100000Stage-II minimum region size in pixels; the final merge removes regions smaller than this.
spatial_weightoptFLOAT0.600–1Weight of spatial distance inside the SLIC clustering (0 = colour/texture only, 1 = pure proximity). The default already leans on colour.
superpixel_sizeoptINT82–256SLIC superpixel size in pixels. Smaller = more seeds = finer detail captured; larger runs faster and over-merges small features.
slic_iterationsoptINT51–50SLIC clustering iterations. 5 is plenty for convergence; raise to 10 on very textured content.

Outputs (3)

NameTypeDescription
segmentedCOMFY_MATCHTYPE_V3The input with every region repainted by its average colour - a quick visual of the region map.
labelsNPARRAYuint16 per-pixel region index (0..num_regions-1), the same height/width as the input. Feed this into mask/crop/Draw Nodes.
num_regionsINTNumber of distinct regions in the label map.