cv2.fastNlMeansDenoising
The classic denoiser, and yes it really is that slow
- src
- result
Non-local means is the pre-deep-learning denoiser that refused to die, and it earns its keep on a specific kind of noise. The idea is worth thirty seconds, because it explains both why it preserves texture where a Gaussian would smear it, and why it takes so long that this pack's source runs it out-of-process.
How it works
Ordinary blurring averages pixels that are near each other. Non-local means averages pixels that look like each other: for each pixel it compares a small patch (templateWindowSize, default 7) against patches throughout a larger search area (searchWindowSize, default 21), and weights each neighbour by how similar its patch is. So a flat region averages heavily over its many similar patches and comes out clean, while an edge only averages along itself and stays sharp. That's why NLM is the tool that keeps detail where a Gaussian flattens it.
The cost lives in that search: work scales linearly with searchWindowSize and with the square of the patch, per pixel. That's why the pack lists fastNlMeansDenoising and its colour sibling among the handful of calls it always executes in an interruptible subprocess - its source notes these take "tens of seconds to minutes on ordinary frames". The upside of that offloading is real: pressing Cancel actually stops the call instead of waiting out a frozen UI.
The inputs that matter
src accepts an IMAGE, MASK or NPARRAY, and the output result echoes the format - IMAGE in, IMAGE out, MASK in, MASK out, raw array stays raw. Three knobs, and only one of them needs thought:
h(default 3.0) - the strength. The tooltip is the honest version: a bighremoves noise and detail, a small one keeps detail and some noise. At 3.0 you get a gentle clean; at 10+ you're into plastic territory on skin. This is the knob you tune.templateWindowSize(7) - patch size. Odd numbers; 7 is the recommended value and there's rarely a reason to leave it.searchWindowSize(21) - how far it looks. Odd, recommended 21, and it's the one with the linear cost. Dropping it to 11 or 15 is the standard "make this bearable on a 4K frame" move, at a small quality cost.
One thing the tooltips make explicit: OpenCV expects this on grayscale and documents the colour path separately, so a multi-channel input is denoised channel by channel with no shared model - fine on sensor noise, occasionally odd on coloured chroma noise. If your input is colour and the noise is chroma, use the Colored variant instead; it converts to CIELAB and handles the colour components properly.
Where it fits
Cleaning an input is the realistic use, not polishing an output. A noisy reference photo used as an img2img base or a composition reference drags its noise into the generation, and a couple of seconds of NLM before that costs less than rerolling. Scanned photos, high-ISO captures, stills from compressed video - those are NLM's home turf. For skin smoothing and AI-grain cleanup, though, the better tools are edge-preserving filters like bilateral or guided, and if the problem is grain, post-processing.md makes the case for why you probably want to be adding some rather than removing it.
For masks, this denoises the boundary of a soft mask nicely - a fuzzy 0–255 edge becomes a clean one - though core's FeatherMask and a small blur usually get you there for free.
Installing it
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"
ComfyUI Manager: search comfyui_cv, install, restart. Python ≥ 3.12 and a recent V3-API ComfyUI. Node path: image/CV/low-level/cv2 F.
Traps
Batching. This wrapper's tooltip states the general contract - an IMAGE is handled as frame 0 of a batch - so if you feed an eight-frame batch, seven frames come back untouched and you will spend twenty minutes wondering why one frame looks different. Break the batch apart or accept that this is a single-image node.
Then the expectations. It is not fast, regardless of the name; the fast refers to the algorithm's improvement over brute-force NLM, not to your wall clock. Feeding it a 4K frame at default settings is a genuine coffee break. And be careful with h on video: tuned per-frame values produce flicker, so pick one and hold it if you're processing a sequence.
Inputs (4)
| Name | Type | Default | Description |
|---|---|---|---|
| src | COMFY_MATCHTYPE_V3 | Input 8-bit 1-channel, 2-channel, 3-channel or 4-channel image. The image output(s) echo this input's format. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| hopt | FLOAT | 3.0000-1e+38–1e+38 | Parameter regulating filter strength. Big h value perfectly removes noise but also removes image details, smaller h value preserves details but also preserves some noise This function expected to be applied to grayscale images. For colored images look at fastNlMeansDenoisingColored. Advanced usage of this functions can be manual denoising of colored image in different colorspaces. Such approach is used in fastNlMeansDenoisingColored by converting image to CIELAB colorspace and then separately denoise L and AB components with different h parameter. Preset to the OpenCV default (3.0). |
| templateWindowSizeopt | INT | 7-2147483648–2147483647 | Size in pixels of the template patch that is used to compute weights. Should be odd. Recommended value 7 pixels Preset to the OpenCV default (7). |
| searchWindowSizeopt | INT | 21-2147483648–2147483647 | Size in pixels of the window that is used to compute weighted average for given pixel. Should be odd. Affect performance linearly: greater searchWindowsSize - greater denoising time. Recommended value 21 pixels Preset to the OpenCV default (21). |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| result | COMFY_MATCHTYPE_V3 | Echoes the 'src' input's format: an IMAGE link comes back as IMAGE, MASK as MASK, NPARRAY stays NPARRAY. |