Nodes/ComfyUI CV/cv2.estimateAffine2D
ComfyUI Node

cv2.estimateAffine2D

Fit a transform to two point sets and get the outliers told on

By bmad4ever·Created 3 months ago·Updated 14 days ago· 1
cv2.estimateAffine2D
  • from_
  • to
  • out
  • inliers
◄methodRANSAC►
◄ransacReprojThreshold3.0000►
◄maxIters2000►
◄confidence0.9900►
◄refineIters10►

This is the node you reach for when you have correspondences and want the transform, rather than wanting to guess it with sliders. Give it matched 2D points from A and the same points in B, and it returns the 2×3 affine matrix that best maps one onto the other - plus a mask telling you which of your matches it decided were garbage. That last output is the part that makes it usable on real feature matches instead of clean, hand-typed point lists.

What it fits, and what it can't

An affine has six degrees of freedom: rotation, non-uniform scale, shear, translation. Three point pairs are theoretically enough; in practice you feed it tens or hundreds. It fits them by least squares in one shot, but the default path is RANSAC, which repeatedly samples minimal subsets, counts inliers within a reprojection threshold, and keeps the consensus - then a Levenberg-Marquardt pass refines that fit. That's why the outputs include inliers: the whole point is that some of your matches are wrong, and this node is designed to survive that.

It cannot do perspective. If the camera moved relative to a planar scene, or a face is turned enough that the near side is noticeably larger, the correct model is a homography - that's CV Find Homography (RANSAC) in this same pack, a curated node with a found flag and an identity fallback. Affine is for the cases where the four-point perspective term is genuinely negligible: small motion, telephoto-ish views, alignment of flat scans, stabilisation between two frames of the same shot.

The inputs

from_ and to are NPARRAY only, both point sets, same length, N x 2, and paired in order - element i of one corresponds to element i of the other. The name from_ has a trailing underscore because from is a Python keyword; that's the only reason it looks odd. In practice these arrive from the pack's curated CV Detect Features → CV Match Features pair, or from CV Points if you're typing a correspondence by hand, or from a tracker's output.

The knobs you'll actually touch:

  • method - RANSAC (default) or LMEDS (least-median, better when more than half your matches are junk and you have no idea what a pixel of error means).
  • ransacReprojThreshold - default 3.0, the maximum reprojection error, in pixels, for a point to count as an inlier. This is the one to think about. Features matched across a resize or an upscale need this in the target image's pixel scale.
  • confidence (0.99) and maxIters (2000) - RANSAC's budget. The tooltip's own advice is good: 0.95–0.99 is plenty, and pushing toward 1 slows the estimate down for no gain, while going under ~0.8–0.9 starts returning wrong transforms.
  • refineIters (10) - the LM refinement. Setting it to 0 disables refining entirely and hands you the raw robust result, which is occasionally what you want when refinement is dragging a good inlier consensus toward a few big outliers.

Outputs are out, an NPARRAY holding the 2×3 matrix, and inliers, the N×1 mask of survivors. out wires straight into cv2.warpAffine - feed the matrix to its M socket, and the size for dsize comes from CV Array Size, which reads the width and height of an ndarray for exactly this reason. warpAffine echoes IMAGE in / IMAGE out, so the whole chain from matches to warped picture stays in normal ComfyUI types.

Where it fits

Alignment before a generative step is the realistic use: straightening a photographed document before OCR-ish processing, aligning two frames so a difference image means something, registering a reference photo to its upscaled or restyled twin. In the identity corner it's the same shape of problem - every face swap pipeline aligns a crop to a canonical template before it does anything else, which is why InsightFace eats most of that market; the pack's CV Affine Shape Warp is the curated version of that idea, and it's the node to look at before you hand-build a warp.

Installing it

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"

Or install comfyui_cv from Manager and restart. Python ≥ 3.12, V3-API ComfyUI. Find it under image/CV/low-level/cv2 E.

Traps

Points must be float, 2-D, and the two sets the same length - mismatched lengths raise rather than truncate. Degenerate inputs (all points collinear, or three points that are nearly one line) give you a bad matrix with no error, so check CV Inspect CV Data and count your inliers before trusting out: a matrix fit from six of four hundred matches is a red flag, not a result. And cross-check the units on ransacReprojThreshold whenever there's an upscale anywhere in the graph - a threshold of 3 pixels means something completely different at 512×512 and 4096×4096.

Categoryimage/CV/low-level/cv2 E

Inputs (7)

NameTypeDefaultDescription
from_NPARRAY - - - A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here.
toNPARRAYSecond input 2D point set containing $(x,y)$. A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here.
methodoptCOMBORANSACRobust method used to compute transformation. The following methods are possible: - - RANSAC-based robust method - - Least-Median robust method RANSAC is the default method.
ransacReprojThresholdoptFLOAT3.0000-1e+38–1e+38Maximum reprojection error in the RANSAC algorithm to consider a point as an inlier. Applies only to RANSAC. Preset to the OpenCV default (3.0).
maxItersoptINT2000-2147483648–2147483647The maximum number of robust method iterations. Preset to the OpenCV default (2000).
confidenceoptFLOAT0.9900-1e+38–1e+38Confidence level, between 0 and 1, for the estimated transformation. Anything between 0.95 and 0.99 is usually good enough. Values too close to 1 can slow down the estimation significantly. Values lower than 0.8-0.9 can result in an incorrectly estimated transformation. Preset to the OpenCV default (0.99).
refineItersoptINT10-2147483648–2147483647Maximum number of iterations of refining algorithm (Levenberg-Marquardt). Passing 0 will disable refining, so the output matrix will be output of robust method. Preset to the OpenCV default (10).

Outputs (2)

NameTypeDescription
outNPARRAY—
inliersNPARRAY—