cv2.findTransformECC (2/2)
Cv2.findTransformECC (2/2)
- templateImage
- inputImage
- warpMatrix
- inputMask
- float
- nparray
Two cv2.findTransformECC nodes exist in this pack because cv2 has two overloads of the function. This is the shorter signature: only the two images and the warp matrix are required, and the wrapper presets OpenCV's documented defaults for the motion model, the stopping criteria and the pre-blur. In practice, that's the one you start with.
ECC aligns two images by maximising the enhanced correlation coefficient between their intensities. No keypoints, no descriptors, no matching - which is why it works on scenes that feature detection can't handle and why it's the standard answer for frame-to-frame registration in a burst, template alignment, and stabilisation of footage that isn't moving much.
Inputs
templateImage is the reference you're aligning to, inputImage is the one that moves. Both accept IMAGE, MASK or NPARRAY; an IMAGE arrives as uint8 BGR frame 0. They need to be the same type - 8-bit, 16-bit, float32 or float64, 1 or 3 channels.
warpMatrix is the part that surprises people: it's a required input. cv2 refines the matrix you hand it in place, so you supply a starting guess and read the refined result off the output. Give it an identity 2×3 float matrix for the default motion models (3×3 for homography). An identity is correct when your frames are close together; for anything with a large displacement or rotation you need a rough prior, because ECC is a local optimiser and will happily converge to a wrong alignment.
motionType defaults to MOTION_AFFINE - 6 degrees of freedom. MOTION_TRANSLATION (2) and MOTION_EUCLIDEAN (3) are stricter and converge more reliably; MOTION_HOMOGRAPHY (8) needs a 3×3 matrix and a good initial guess. Pick the smallest model that describes the motion you actually have.
criteria_type / criteria_max_count / criteria_epsilon are the split TermCriteria: stop after n iterations, stop when the correlation stops improving by epsilon, or whichever comes first. The presets are 30 and 0.001 - cv2's own defaults, and conservative ones. For real registration work, raise the iteration count (the shipped ECC exercise runs 200 with epsilon 1e-6, and the curated node defaults to 100 at 1e-5). Thirty iterations is often where you find out the answer was "almost aligned".
inputMask is optional, single channel, and marks which pixels of inputImage count. Letterbox bars, watermarks, a scratch - mask them and the optimiser stops trying to explain them.
Note what's missing relative to the (1/2) overload: there's no gaussFiltSize widget here at all. OpenCV's internal default (an odd 5) applies, which is also why this overload is less of a minefield - the (1/2) node exposes that parameter with a widget default of 0 while the documentation says 5.
Outputs
float is the final correlation coefficient, where 1 is a perfect match - your quality signal, and worth checking before you trust the warp. nparray is the refined transform.
Applying it: cv2_warpAffine (or cv2_warpPerspective for homography) with INTER_LINEAR | WARP_INVERSE_MAP in the flags and dsize set to the template's size. The inverse-map flag is not optional here - the matrix maps template coordinates into the input, so you warp the input backwards through it to land in the template's frame. exercise_image_registration_ecc.json builds the whole thing, including the denoising payoff: average three aligned noisy frames and the grain drops; average them unaligned and you get ghosting.
Failure mode, and the friendlier alternative
cv2 throws when the iteration doesn't converge - a featureless frame, two unrelated images, zero variance. In a graph, that exception ends the run. The curated CV Find Transform (ECC) node wraps the same function with a found flag and an identity fallback, plus a min_correlation threshold you can set to reject weak alignments outright. If you're feeding it frames you haven't vetted, use that node. If you're feeding it frames you control and want the exact parameters, this wrapper is fine.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"
ComfyUI Manager: search "ComfyUI CV", install, restart. Python ≥ 3.12, recent V3-API ComfyUI.
Common issues
Workflow dies with an OpenCV error. Non-convergence. Downscale, grayscale, add a pre-blur, or start from a rough warp rather than identity.
Correlation comes out around 0.5–0.7 and the image is still misaligned. You're in a local minimum, typically because the starting matrix was too far off. Or your motion model is wrong for the content - a parallax-heavy scene has no single affine answer.
It's slow. It's iterative and runs on the CPU. Align on downscaled copies and apply the resulting matrix to the full-resolution image; the geometry scales fine, the accuracy mostly survives.
Inputs (8)
| Name | Type | Default | Description |
|---|---|---|---|
| templateImage | NPARRAY,IMAGE,MASK | 1 or 3 channel template image; CV_8U, CV_16U, CV_32F, CV_64F type. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| inputImage | NPARRAY,IMAGE,MASK | input image which should be warped with the final warpMatrix in order to provide an image similar to templateImage, same type as templateImage. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| warpMatrix | NPARRAY | floating-point $2\times 3$ or $3\times 3$ mapping matrix (warp). A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here. | |
| motionTypeopt | COMBO | MOTION_AFFINE | parameter, specifying the type of motion: - **MOTION_TRANSLATION** sets a translational motion model; warpMatrix is $2\times 3$ with the first $2\times 2$ part being the unity matrix and the rest two parameters being estimated. - **MOTION_EUCLIDEAN** sets a Euclidean (rigid) transformation as motion model; three parameters are estimated; warpMatrix is $2\times 3$. - **MOTION_AFFINE** sets an affine motion model (DEFAULT); six parameters are estimated; warpMatrix is $2\times 3$. - **MOTION_HOMOGRAPHY** sets a homography as a motion model; eight parameters are estimated;\`warpMatrix\` is $3\times 3$. |
| criteria_typeopt | COMBO | max count or epsilon (whichever first) | When to stop iterating: after max_count iterations, when the change drops below epsilon, or whichever comes first. |
| criteria_max_countopt | INT | 301–2147483647 | Maximum iterations (ignored when 'epsilon only'). |
| criteria_epsilonopt | FLOAT | 0.000–1e+38 | Target accuracy / smallest change worth continuing for (ignored when 'max count only'). |
| inputMaskopt | NPARRAY,IMAGE,MASK | An optional single channel mask to indicate valid values of inputImage. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| float | FLOAT | — |
| nparray | NPARRAY | — |