CV Match Image Moments
Shape matching that can tell two shaded discs apart
- reference
- images
- distances
- matched
- best_index
- found
- scale_ratios
- rotations
- mirrored
- transforms
Contour matching has a blind spot, and it's a big one. Give it two identical outlines and it reports distance 0.000000 - because it doesn't see inside them. Two discs lit from opposite sides are the same object to matchShapes on a contour, and completely different objects to a human. If the thing distinguishing your two candidates is internal shading - a lit face versus a shadowed one, a chip versus a blank, a printed label versus a plain surface - contour matching will confidently tell you they're identical.
This node scores frames by Hu moments of the image intensity instead, using the same scale/rotation/mirror invariance policy as the pack's contour matcher (CV Filter Contours By Shape). cv2.moments on a raster defaults to binaryImage=False, meaning a pixel's mass is its value - Hu's original formulation. That's the whole idea: shading is geometry.
How it works, and where it gets slippery
reference is your template (frame 0 if you hand it a batch); images is the batch to score. Size doesn't have to match - moments are scale invariant.
Three things you actually tune:
weighting.intensity, peak-normalizedis the default and the right starting point, and the reason is concrete: Hu moments are not invariant to a brightness change, so a half-exposure copy of the same object scores differently. Dividing by the frame's own maximum more than halved that error in the author's own testing (I1 0.0394 raw → 0.0146). It doesn't fully remove it - dimming also shrinks the effective support of a soft-edged region and no normalization fixes that.intensityis the raw formulation.binarythrows away values and uses the silhouette, which is basically what the contour nodes do - with a nasty edge: every non-zero pixel counts as shape, so any background haze swallows the entire frame.max_distance(default 0.3). "0 = identical". Intensity distances do not share a scale with contour distances, so calibrate: crank it to 1e6, read thedistancesoutput on real data, then set a threshold that means something.method-CONTOURS_MATCH_I1, I2, I3. I1 is the usual choice.
Then the invariance policy: scale_mode / scale_tolerance (note that on this path "size" is sqrt(total intensity), so a brighter copy reads as bigger - unless weighting is binary), rotation_mode (the "or 180 flipped" option exists for 2-fold symmetric shapes whose orientation flips arbitrarily), rotation_tolerance_deg, mirror_mode, and min_chirality.
That last one is worth understanding rather than skipping. Handedness is decided by the sign of the 7th Hu invariant - the one OpenCV's own docs note is only proved invariant "with the assumption of infinite image resolution." Rasterize a symmetric shape and you get a small non-zero h7 whose sign is noise. min_chirality (default 1e-6) is the magnitude below which mirroring is reported as undecidable instead of guessed. Hard-edged silhouettes measure ~0.005 of noise; a smooth intensity field measures ~1e-5, which is why this default is far lower than the contour matcher's.
Inputs and outputs
reference, images (both accept IMAGE/MASK or NPARRAY directly), method, max_distance, weighting, plus the four invariance knobs.
Outputs: distances (N, float32, in batch order - deliberately not sorted, because batch order is what CV Index Batch needs), matched (N, float32 keep-mask: 1 where the frame passed both the threshold and the policy), best_index (closest accepted frame, or −1), found (boolean - wire it into an if/else rather than testing best_index by hand), scale_ratios, rotations, mirrored, and transforms - an (N,2,3) similarity matrix per frame mapping reference coordinates onto that frame, which you can hand straight to cv2_warpAffine to overlay your template on its match.
The keep-mask shape is deliberate, not lazy: it means a zero-match result stays a valid batch, so downstream batch math doesn't get a ragged input.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
Or Manager → ComfyUI CV. Restart and reload. Pure cv2, no models, one dependency (opencv-contrib-python-headless~=5.0.0.93), Python ≥ 3.12, V3-API-era ComfyUI.
Common issues
Distances that look tiny and frames that still don't match. You're probably comparing an intensity distance against a contour-shaped expectation. Read distances and calibrate max_distance.
A background haze matching everything. You set weighting to binary on an image that isn't a clean silhouette. That mode treats any non-zero pixel as part of the shape.
Rotation output that looks like noise on a square. It is noise - a square has no well-defined orientation, and the node's own tooltip says so. Check pose_confidence on CV Image Moments before trusting orientation-based filtering.
Inputs (11)
| Name | Type | Default | Description |
|---|---|---|---|
| reference | NPARRAY,IMAGE,MASK | The template to match against (frame 0 if a batch). It does NOT have to be the same size as the candidates - moments are scale invariant. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| images | NPARRAY,IMAGE,MASK | Candidate frames, one score per frame. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| method | COMBO | CONTOURS_MATCH_I1 | How the Hu signatures are compared (I1 is the usual choice; I2/I3 are alternative norms). |
| max_distance | FLOAT | 0.300–1000000 | Largest accepted dissimilarity (0 = identical). Raise it to 1e6 and read the distances output to calibrate - intensity distances do not have the same scale as contour ones. |
| weighting | COMBO | intensity, peak-normalized | Which density function the moments integrate. 'intensity' is Hu's original formulation and cv2's default (mass = pixel value), so a shaded region is a different shape from a flat one. Hu moments are NOT invariant to a brightness gain, so 'peak-normalized' (divide by the frame's own maximum) is the default: it more than halved the error of a x0.5 exposure change in testing (I1 0.0394 raw -> 0.0146). It does NOT fully remove the effect - dimming also shrinks the effective support of a soft-edged region, and no normalization can fix that. 'binary' ignores the values and uses the silhouette, which is the closest thing to what the contour nodes measure - but beware, it treats EVERY non-zero pixel as part of the shape, so any background haze swallows the whole frame. |
| scale_modeopt | COMBO | any scale (invariant) | Hu moments ignore size. 'same scale as reference' additionally requires the candidate to be about as big as the reference - the ratio of sqrt(m00), so 0.5 means half the linear size. NOTE: on this path 'size' is sqrt(total intensity), so a brighter copy reads as a bigger one unless weighting is 'binary'. |
| scale_toleranceopt | FLOAT | 0.100–100 | Accepted size band as a fraction: 0.10 accepts 0.91x to 1.10x of the reference (symmetric in the ratio, so growing and shrinking are treated alike). 0 demands an exact match. Ignored unless scale is 'same scale as reference'. |
| rotation_modeopt | COMBO | any rotation (invariant) | 'same orientation as reference' keeps only candidates standing the same way up. 'same orientation or 180 deg flipped' also accepts the upside-down copy - USE IT for 2-fold symmetric shapes (ellipse, rectangle), whose 0-360 orientation flips arbitrarily. Both are meaningless for a shape with no well-defined orientation: check 'pose_confidence' first (a rotating square reads 0, 140, 145, 155, 135 degrees - pure noise). |
| rotation_tolerance_degopt | FLOAT | 100–180 | Half-width of the accepted rotation window in degrees. Compared CIRCULARLY, so 359 degrees counts as 1 degree away from 0. |
| mirror_modeopt | COMBO | either handedness (invariant) | The matchShapes distance is a poor handedness test, so decide it here instead. Only the 7th Hu invariant's SIGN changes under reflection (OpenCV: "invariants to the image scale, rotation, and reflection except the seventh one, whose sign is changed by reflection"), and matchShapes skips any term below its internal eps = 1e-5 - which |h7| usually is. Measured: an 'F' and its mirror are indistinguishable (~0), while a hook with a bigger |h7| scores 0.50. 'same handedness only' drops reflected copies; 'mirrored only' keeps just the reflected ones. A mirror-symmetric shape (square, circle, isoceles triangle) is its own mirror, so it counts as UNDECIDED: kept by 'same handedness only', dropped by 'mirrored only'. |
| min_chiralityopt | FLOAT | 0.00000–1 | How chiral a shape must be before mirroring can be decided at all - the magnitude of the 'chirality' value. This knob exists because the 7th Hu invariant is the one routinely discarded in practice: it is the smallest and the most fragile, and OpenCV's own docs note the invariance is proved "with the assumption of infinite image resolution", so "in case of raster images, the computed Hu invariants for the original and transformed images are a bit different". Rasterizing a symmetric shape therefore produces a small NON-zero h7 whose sign is meaningless. Measured magnitudes: an 'F' glyph 0.020-0.048, a scalene triangle 0.0069, a rasterized symmetric shape up to ~0.005 of pure noise, a square exactly 0. Below this on EITHER side the match is reported as mirrored = 0 (undecidable) rather than guessed. Set 0 to trust the sign always. Those numbers are for hard-edged silhouettes; a smooth intensity field measures ~1e-5, which is why the default here is far lower than on the contour matcher. Read the 'chirality' output of 'CV Image Moments' on your own data before raising it. |
Outputs (8)
| Name | Type | Description |
|---|---|---|
| distances | NPARRAY | (N,) float32 matchShapes distance per INPUT frame, in batch order (not sorted - the batch order is what 'CV Index Batch' needs). Rejected frames keep their real distance; use 'matched' to tell them apart. |
| matched | NPARRAY | (N,) float32 keep-mask, 1 where the frame passed BOTH the distance threshold and the invariance policy, else 0. |
| best_index | INT | Index of the closest ACCEPTED frame, or -1 when nothing matched. Feed it to 'CV Index Batch'. |
| found | BOOLEAN | False when no frame passed - wire it into an 'if/else' rather than testing best_index by hand. |
| scale_ratios | NPARRAY | (N,) float32 sqrt(m00) of each frame over the reference's. 0.0 where no pose exists. |
| rotations | NPARRAY | (N,) float32 rotation in degrees, 0-360, taking the reference onto that frame (Y down, clockwise on screen). For a MIRRORED match it is the rotation applied AFTER flipping the reference about its vertical axis. |
| mirrored | NPARRAY | (N,) float32 handedness verdict: +1 reflected, -1 same handedness, 0 undecidable (symmetric, or below 'min_chirality'). |
| transforms | NPARRAY | (N,2,3) float32 similarity matrix per frame mapping REFERENCE coordinates onto that frame (mirror, scale, rotate, centroid to centroid). Feed a row to cv2_warpAffine to overlay the template on its match. |