Nodes/ComfyUI CV/cv2.findTransformECC (1/2)
ComfyUI Node

cv2.findTransformECC (1/2)

Cv2.findTransformECC (1/2)

By bmad4ever·Created 4 months ago·Updated 14 days ago· 1
cv2.findTransformECC (1/2)
  • templateImage
  • inputImage
  • warpMatrix
  • inputMask
  • float
  • nparray
◄motionTypeMOTION_AFFINE►
◄criteria_typemax count or epsilon (whichever first)►
◄criteria_max_count30►
◄criteria_epsilon0.00►
◄gaussFiltSize0►

ECC - enhanced correlation coefficient - is the other way to align two images. No keypoints, no matching, no RANSAC. It optimises a warp so that one image's intensities correlate with the other's as strongly as possible. That makes it the right tool when your two frames are the same scene, slightly moved: frames of a burst, a golden reference template, a photo pair shot from a tripod, a document that got nudged between scans.

This node is the full-parameter overload of cv2.findTransformECC. Every knob cv2 has is here, and for once that's not a misleading boast: motionType, the termination criteria, the mask and the pre-blur size are all exposed rather than hidden.

The odd one out: warpMatrix is an input

Most estimation nodes return a matrix. This one takes one and returns it. warpMatrix is a required NPARRAY input - the initial guess - which cv2 refines in place, and the refined result comes back as the node's second output. With no better information you feed it an identity: for MOTION_TRANSLATION, MOTION_EUCLIDEAN and MOTION_AFFINE that's a 2×3 float matrix; for MOTION_HOMOGRAPHY it must be 3×3. (cv2 ignores the third row of a 3×3 when the motion type is one of the first three.)

This is also where the honest limitation lives: ECC is a local optimiser. If the frames are a couple of degrees and a handful of pixels apart, identity is a fine starting point. If they're badly displaced or rotated, you need a rough alignment first - a similarity transform from a cheaper estimator, or the previous frame's result in a tracking loop - or the optimiser marches off toward a wrong local maximum and reports it with confidence.

Inputs that matter

templateImage and inputImage are the two frames: 1- or 3-channel, CV_8U/16U/32F/64F, mutually the same. The sockets are image-ish, so IMAGE, MASK or NPARRAY all work - an IMAGE arrives as uint8 BGR frame 0, which is the normal case for a ComfyUI graph.

motionType sets the model: MOTION_TRANSLATION (2 degrees of freedom), MOTION_EUCLIDEAN (shift plus rotation, 3), MOTION_AFFINE (6, the default) or MOTION_HOMOGRAPHY (8). Fewer degrees of freedom converge more reliably, and the failure mode of too many is a warp that fits noise. If your images differ only by a shift, don't ask for a homography.

criteria_type / criteria_max_count / criteria_epsilon are one cv2 TermCriteria split into three widgets - stop after n iterations, stop when the improvement drops below epsilon, or whichever comes first. The defaults are 30 iterations and 0.001, which are conservative: a real alignment usually wants more iterations and a tighter epsilon.

inputMask is a single-channel mask saying which pixels of inputImage are valid. Useful for letterboxed frames, watermarks and sensor artefacts.

gaussFiltSize deserves a warning. The widget defaults to 0, but OpenCV documents 5 as the default and this parameter wants an odd size - it's a pre-blur that stops the optimiser chasing sensor noise. Set 5, or another odd number, rather than trusting the default. The curated CV Find Transform (ECC) node forces an odd value for you, which is one of several reasons it's often the better pick.

Outputs

float is the final ECC correlation coefficient - cv2's return value, 1 meaning a perfect match. It's your quality score, and it's the number to log and branch on. nparray is the refined warp matrix, ready for cv2_warpAffine. Apply it with WARP_INVERSE_MAP in the flags and dsize set to the template's size, which is what pulls the input into the template's frame - the shipped exercise_image_registration_ecc.json does exactly that, and shows the payoff (aligned stack average, sharp; unaligned, ghosted).

The failure mode to know about

OpenCV's own documentation is explicit: the function throws when the algorithm doesn't converge - featureless frames, unrelated images, degenerate contrast. In a ComfyUI graph a cv2 exception stops the run. If you're processing a folder where some frames are junk, the curated CV Find Transform (ECC) node exists precisely because of this: it catches the non-convergence and returns found = false with an identity matrix, so a control-flow node can skip the bad frames instead of killing the queue.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"

Manager users: "ComfyUI CV", install, restart. Python ≥ 3.12 and a V3-API ComfyUI.

Common issues

OpenCV error, workflow dead. Non-convergence, as above. Shrink the problem: grayscale the frames, downscale before aligning and scale the resulting matrix back up, or start from a rough warp instead of identity.

It converges to a wrong answer. ECC is iterative and greedy; with big displacements it lands in a local optimum. Give it a starting matrix that's roughly right.

Deceptively slow runs. Whole-frame alignment on 4K pairs with 500 iterations is not fast, and it's CPU work. Resize first, refine second.

Categoryimage/CV/low-level/cv2 F

Inputs (9)

NameTypeDefaultDescription
templateImageNPARRAY,IMAGE,MASK1 or 3 channel template image; CV_8U, CV_16U, CV_32F, CV_64F type. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
inputImageNPARRAY,IMAGE,MASKinput image which should be warped with the final warpMatrix in order to provide an image similar to templateImage, same type as templateImage. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
warpMatrixNPARRAYfloating-point $2\times 3$ or $3\times 3$ mapping matrix (warp). A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here.
motionTypeCOMBOMOTION_AFFINEparameter, specifying the type of motion: - **MOTION_TRANSLATION** sets a translational motion model; warpMatrix is $2\times 3$ with the first $2\times 2$ part being the unity matrix and the rest two parameters being estimated. - **MOTION_EUCLIDEAN** sets a Euclidean (rigid) transformation as motion model; three parameters are estimated; warpMatrix is $2\times 3$. - **MOTION_AFFINE** sets an affine motion model (DEFAULT); six parameters are estimated; warpMatrix is $2\times 3$. - **MOTION_HOMOGRAPHY** sets a homography as a motion model; eight parameters are estimated;\`warpMatrix\` is $3\times 3$.
criteria_typeCOMBOmax count or epsilon (whichever first)When to stop iterating: after max_count iterations, when the change drops below epsilon, or whichever comes first.
criteria_max_countINT301–2147483647Maximum iterations (ignored when 'epsilon only').
criteria_epsilonFLOAT0.000–1e+38Target accuracy / smallest change worth continuing for (ignored when 'max count only').
inputMaskNPARRAY,IMAGE,MASKAn optional single channel mask to indicate valid values of inputImage. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
gaussFiltSizeINT0-2147483648–2147483647An optional value indicating size of gaussian blur filter; (DEFAULT: 5) The function estimates the optimum transformation (warpMatrix) with respect to ECC criterion (), that is $$\texttt{warpMatrix} = \arg\max_{W} \texttt{ECC}(\texttt{templateImage}(x,y),\texttt{inputImage}(x',y'))$$ where $$\begin{bmatrix} x' \ y' \end{bmatrix} = W \cdot \begin{bmatrix} x \ y \ 1 \end{bmatrix}$$ (the equation holds with homogeneous coordinates for homography). It returns the final enhanced correlation coefficient, that is the correlation coefficient between the template image and the final warped input image. When a $3\times 3$ matrix is given with motionType =0, 1 or 2, the third row is ignored. Unlike findHomography and estimateRigidTransform, the function findTransformECC implements an area-based alignment that builds on intensity similarities. In essence, the function updates the initial transformation that roughly aligns the images. If this information is missing, the identity warp (unity matrix) is used as an initialization. Note that if images undergo strong displacements/rotations, an initial transformation that roughly aligns the images is necessary (e.g., a simple euclidean/similarity transform that allows for the images showing the same image content approximately). Use inverse warping in the second image to take an image close to the first one, i.e. use the flag WARP_INVERSE_MAP with warpAffine or warpPerspective. See also the OpenCV sample image_alignment.cpp that demonstrates the use of the function. Note that the function throws an exception if algorithm does not converges.

Outputs (2)

NameTypeDescription
floatFLOAT—
nparrayNPARRAY—