CV Find Transform (ECC)
Align two frames by intensity, no keypoints required
- template
- image
- mask
- warp_matrix
- correlation
- found
Feature matching needs features. A blurry burst frame, a star field, a flat wall, a clean plate with nothing distinctive in it - SIFT and ORB come back nearly empty, and your alignment dies. This node aligns by intensity instead: it maximizes the Enhanced Correlation Coefficient between two images and returns the transform, no keypoints, no descriptors, no matcher.
It's cv2.findTransformECC from bmad4ever's ComfyUI CV pack (a fork of Gerold Meisinger's opencv-comfyui, ~470 raw cv2.* wrappers plus curated nodes). Deterministic, CPU-cheap, and the natural partner to the pack's MTB alignment nodes for HDR brackets and to image-difference work like clean-plate matting.
How it works
Call it the "nudge until it lines up" method. Both images are converted to grayscale, optionally blurred by gauss_filt_size (default 5, odd, 1 = no smoothing - the blur exists so the optimizer doesn't chase sensor noise), and then the node searches for the transform that maximizes correlation between them.
motion_type is the shape of what you're searching for, and it's the setting that decides whether this works:
MOTION_TRANSLATION- shift only, 2 degrees of freedom.MOTION_EUCLIDEAN- shift plus rotation, 3 dof. The default, and the right one for a hand-held burst.MOTION_AFFINE- 6 dof.MOTION_HOMOGRAPHY- 8 dof, output is 3×3 instead of 2×3.
Fewer degrees of freedom converge more reliably. Ask for a homography on a pair that only shifted and you're inviting a bad local minimum.
iterations (100) and epsilon (1e-5) are the stop conditions: it quits when the correlation stops improving by more than epsilon, or after iterations. min_correlation (0.0) is the honesty knob - ECC returns a coefficient where 1.0 is a perfect match, and found is only true if the final value beats your threshold. Leave it at 0 and any converged estimate counts as found; set 0.6–0.9 for anything you're going to composite.
There's an optional mask too, and it means what you'd hope: non-zero pixels are the ones that participate. The pack's docs are blunt that it must be a coherent region - a scattered or speckled mask gives the optimizer nothing to lock onto.
Outputs and how to apply them
warp_matrix is a 2×3 float32 (3×3 for homography) that maps template coordinates into image - i.e. the inverse of the direction you probably assumed. That's why the way you apply it is cv2_warpAffine with its size argument set to the template's dimensions and flags INTER_LINEAR | WARP_INVERSE_MAP, which pulls image into the template's frame. Use CV Warp Flags to author that flag string instead of guessing bit values, and CV Array Size to get the template's width/height as integers.
correlation is the final coefficient - print it with an inspect node while you tune. found is the branch: the raw cv2 call raises when the iteration doesn't converge, which would abort your whole run, and this node catches exactly that and returns found = false with an identity matrix that warps to a no-op. So a bad pair costs you a pass-through, not a failed queue.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"
Or find ComfyUI CV in ComfyUI Manager. Needs Python ≥ 3.12 and a recent, V3-API ComfyUI. To see it working with bundled samples, run workflows/01_install_example_inputs.json once and reload the page before opening workflows/exercise_image_registration_ecc.json.
Common issues
found = false on two obviously related images. ECC is a local optimizer - it needs a starting point close to the answer, and it fails on large displacement or big rotation. It's not a drop-in replacement for the feature-matching path. Crop to roughly the overlapping region first, or start with MOTION_TRANSLATION to get a rough shift, then re-run with a richer model.
Correlation is high but the result is edge-of-frame mush. Featureless or repetitive textures (a wall, a grid) can correlate confidently in the wrong place. gauss_filt_size higher helps a bit; a mask over the distinctive part of the frame helps more.
Grayscale inputs only, effectively. Colour gets converted internally, so don't expect this to fix a colour shift - that's a different job, and the KB's standing advice applies: match colour with a statistics transfer, not a registration node and definitely not a re-generation.
Nothing registers in the node list. The pack needs the V3 node API. On an older ComfyUI you get an import failure in the console, not a broken node.
Inputs (8)
| Name | Type | Default | Description |
|---|---|---|---|
| template | NPARRAY,IMAGE | Reference image the other one is aligned TO. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| image | NPARRAY,IMAGE | Image to align onto the template. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| motion_type | COMBO | MOTION_EUCLIDEAN | Transform model to estimate: TRANSLATION (2 dof), EUCLIDEAN (shift + rotation, 3 dof), AFFINE (6 dof) or HOMOGRAPHY (8 dof, 3x3). Fewer degrees of freedom converge more reliably. |
| iterations | INT | 1001–5000 | Maximum ECC refinement iterations. |
| epsilon | FLOAT | 0.00000–1 | Stop when the correlation improves by less than this between iterations (smaller = more precise, slower). |
| gauss_filt_size | INT | 51–101 | Odd size of the Gaussian blur applied internally before matching - smooths noise so the optimization does not chase it. 1 = no smoothing. |
| min_correlation | FLOAT | 0.00-1–1 | Minimum final ECC correlation for 'found' to be true (1 = perfect). 0 accepts any converged estimate; raise it to reject dubious alignments. |
| maskopt | NPARRAY,MASK | Optional mask on the IMAGE being aligned: non-zero pixels participate in the alignment. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| warp_matrix | NPARRAY | 2x3 float32 transform (3x3 for MOTION_HOMOGRAPHY) mapping template coordinates into the image - use WARP_INVERSE_MAP when warping the image back. Identity if not found. |
| correlation | FLOAT | Final ECC correlation coefficient (1 = perfect alignment; 0.0 when not found). |
| found | BOOLEAN | — |