cv2.pyrMeanShiftFiltering
The flatten-everything filter, and the one input that makes it slow
- src
- result
Mean-shift filtering is the "make this look like a painting" operation, and it's a genuinely useful preprocessing step for segmentation: it flattens colour regions while keeping their boundaries, so a subsequent threshold, region-grow or clustering pass has far less to chew on. The pyramid part means it does this at multiple scales, which is what gives it that characteristic smooth-but-banded look instead of just blurring.
What it does
For every pixel, it looks at the pixels within a spatial radius and keeps only those whose colour is within a second radius, then moves the pixel toward the mean of that set - iterating until things settle, and doing so at several pyramid levels. Two radii, one for space and one for colour, and that's the entire personality of the node:
sp- the spatial window radius. How big a neighbourhood counts as "nearby".10is a sensible start; higher flattens over larger structures.sr- the colour window radius. How different a colour can be and still be merged. This is the one to tune for the look: ~20–40 keeps shading, ~60–80 posterises hard.
maxLevel (default 3) sets how deep the pyramid goes; the term criteria pair (termcrit_type, plus termcrit_max_count and termcrit_epsilon, both visible under advanced inputs) controls when the iteration stops.
The cost is real. Mean shift is not a convolution you can shortcut, and a large sp on a 4K frame is the slowest thing you can do with this pack outside of a DNN node. Downscale, filter, upscale the result if you're after the look rather than the pixels - the pack's own cartoonify example makes the same argument about a large bilateral radius.
How it's wired
src wants a 3-channel 8-bit image - a MASK link is not accepted here, unlike most of the pack's image sockets, because the function is colour-semantic by definition. Optional maxLevel and the term-criteria group are the only other knobs; there's no dstsize, no border type.
The output is result and it echoes the input format: IMAGE in, IMAGE out, so you can drop a preview straight on it. It is not in the pack's per-frame batch set, which means a batch collapses to frame 0 - for video, loop it yourself or use the curated high-level nodes, which handle batches.
Where it earns its place in a generative workflow: as the flattening step before an edge-mask or a segmentation. Flatten, then take edges (cv2.Canny, or an adaptive threshold for ink lines), then use the result as a mask or a stylisation layer. The pack's CV Cartoonify subgraph is the reference version of this idea - it uses a bilateral filter plus an adaptive-threshold ink pass rather than mean shift, and its docs note that a very large bilateral radius is slow on big images, which is the same trade you're making here.
Install
From ComfyUI CV by bmad4ever. ComfyUI Manager → search comfyui_cv, or:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"
Restart. Python ≥ 3.12 and a recent ComfyUI on the V3 node API. No models.
Common issues
"Takes forever / ComfyUI looks hung." It isn't hung, it's mean shift on a big image with a big sp. Start small (sp=10, sr=30), and check the resolution you're feeding it - 2048px on sp=20 is minutes, not seconds.
"Only the first frame changed." Known and by design for this raw wrapper: it isn't batch-aware. Work the batch with a loop, or use the pack's curated image nodes, which are written to handle IMAGE batches.
OpenCV error about channels. You're feeding it a mask or a grayscale NPARRAY. Convert to 3-channel first (cv2.cvtColor with a GRAY2BGR-style code) or feed the original RGB frame.
The result looks muddy, not stylish. sr too high - you've merged shading bands that were holding the shape together. Drop it toward 20 and re-run.
Environment-level weirdness after installing. The usual: OpenCV 5 wants numpy 2.x and packs pinned to numpy 1.x will fight it. Also worth knowing this pack needs Python ≥ 3.12, so an older embedded environment will simply refuse to load it.
Inputs (7)
| Name | Type | Default | Description |
|---|---|---|---|
| src | COMFY_MATCHTYPE_V3 | The source 8-bit, 3-channel image. The image output(s) echo this input's format. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| sp | FLOAT | 0.0000-1e+38–1e+38 | The spatial window radius. |
| sr | FLOAT | 0.0000-1e+38–1e+38 | The color window radius. |
| maxLevelopt | INT | 3-2147483648–2147483647 | Maximum level of the pyramid for the segmentation. Preset to the OpenCV default (3). |
| termcrit_typeopt | COMBO | max count or epsilon (whichever first) | When to stop iterating: after max_count iterations, when the change drops below epsilon, or whichever comes first. |
| termcrit_max_countopt | INT | 301–2147483647 | Maximum iterations (ignored when 'epsilon only'). |
| termcrit_epsilonopt | FLOAT | 0.000–1e+38 | Target accuracy / smallest change worth continuing for (ignored when 'max count only'). |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| result | COMFY_MATCHTYPE_V3 | Echoes the 'src' input's format: an IMAGE link comes back as IMAGE, MASK as MASK, NPARRAY stays NPARRAY. |