Nodes/ComfyUI CV/cv2.findHomography
ComfyUI Node

cv2.findHomography

Cv2.findHomography

By bmad4ever·Created 3 months ago·Updated 14 days ago· 1
cv2.findHomography
  • srcPoints
  • dstPoints
  • mask
  • H
  • mask
◄methodRANSAC►
◄ransacReprojThreshold3.0000►
◄maxIters2000►
◄confidence0.9950►

A homography is the transform between two views of a flat thing. Poster, painting, book cover, screen, wall, floor tile - anything planar. Find four or more point correspondences between the source and the destination, and OpenCV gives you the 3×3 matrix that maps one plane onto the other. That matrix is what puts a rendering into a photograph, snaps a snapped photo back to a rectangle, or aligns two frames of a document.

In this pack, cv2_findHomography is the raw wrapper. It's also the node I'd only pick when I want its exact parameters, because the curated CV Find Homography (RANSAC) next to it adds the two things the raw version lacks: a found flag and a safe identity fallback.

Inputs

srcPoints and dstPoints are the two point sets - typically from CV Match Features for an automatic estimate, or from CV Points / CV Find Contours when you know the corners by hand. Both are NPARRAY-only data sockets (the pack says so in the tooltips; these are coordinate arrays, not pictures) and both must have the same length. cv2 wants floating-point CV_32FC2 data; integer arrays are the usual cause of weird results.

method is the robust estimator: RANSAC (default; handles almost any outlier ratio but needs a threshold), LMEDS (no threshold, but only when more than half your points are inliers), RHO (PROSAC-ish), the USAC_* variants, or least-squares (all points) for the no-outlier case where you just want the plain fit. For feature matches: RANSAC, no debate.

ransacReprojThreshold (3.0 by default) is the maximum reprojection error in pixels for a pair to count as an inlier - OpenCV's own docs suggest 1 to 10 for pixel coordinates. maxIters (2000) and confidence (0.995) are window dressing until the inlier ratio is bad.

That leaves the optional mask input, and the tooltip is unusually blunt about it: input mask values are ignored. It's the C++ output-buffer parameter. The inlier mask is an output.

Outputs

H is the 3×3 perspective transform (determined up to scale, normalized so h33 = 1 when it's non-zero). mask is the Nx1 uint8 inlier mask. Feed the mask to CV Draw Matches and you'll see instantly whether you're looking at a real planar alignment or four points that happened to be near each other.

Two consumers matter. cv2_warpPerspective warps the whole image through H (dsize is a Size - author it with CV Tuple), and cv2_perspectiveTransform pushes individual points through it, which is how you overlay vector geometry or a synthetic layer. The pack's AR examples - exercise_homography_ar.json, exercise_ar_planar_cube.json, exercise_ar_planar_video.json, exercise_multiposter_detection.json - are all built on that pair, and they're the fastest way to see what a good H buys you.

For the case where you already have four corners and don't need robust estimation, cv2_getPerspectiveTransform is the direct four-point solution: no RANSAC, no mask, no surprises. And if you need the opposite direction - "given this H, what planar motion does it mean?" - CV Decompose Homography returns the candidate rotations and translations with a disambiguation step.

The raw-wrapper trade

Here's the thing about using this node rather than the curated one. When findHomography can't find a sensible transform - three collinear points, a degenerate configuration, fewer than four pairs - cv2 can return an empty matrix or a matrix full of garbage, and the raw wrapper hands it to you without comment. Everything downstream then warps an image to nowhere. The curated CV Find Homography (RANSAC) node checks for exactly this: too few points, a non-finite result, or fewer than your minimum inlier count returns an identity homography, a zero mask and found = false, so a control-flow node can skip the compositing. That contract is worth more than parameter access, most days.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"

ComfyUI Manager: search "ComfyUI CV", install, restart. Python ≥ 3.12 and a recent, V3-node-API ComfyUI.

Common issues

H is a mess and the warp looks like a fun-house mirror. Fewer than four non-collinear pairs, or correspondences that aren't really on one plane. If your "plane" has depth in it (a building facade seen at an angle, a person), no homography exists and you want an essential matrix instead.

cv2 error about point types. Cast to float and reshape to Nx1x2 (CV Cast Array, then CV Reshape Array) before feeding the node. Contour output in particular arrives in shapes cv2's homography won't take as-is.

Warped output is black in places. H maps source pixels outside the destination canvas. Check the warp's dsize and border mode, not the homography.

A contributing wheel fight. If nodes from this pack's contrib-dependent areas vanish, a non-contrib OpenCV wheel has likely emptied the shared site-packages/cv2. tools/repair_opencv_contrib.py --check then --apply.

Categoryimage/CV/low-level/cv2 F

Inputs (7)

NameTypeDefaultDescription
srcPointsNPARRAYCoordinates of the points in the original plane, a matrix of the type CV_32FC2 or vector\ . A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here.
dstPointsNPARRAYCoordinates of the points in the target plane, a matrix of the type CV_32FC2 or a vector\ . A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here.
methodoptCOMBORANSACMethod used to compute a homography matrix. The following methods are possible: - **0** - a regular method using all the points, i.e., the least squares method - - RANSAC-based robust method - - Least-Median robust method - - PROSAC-based robust method
ransacReprojThresholdoptFLOAT3.0000-1e+38–1e+38Maximum allowed reprojection error to treat a point pair as an inlier (used in the RANSAC and RHO methods only). That is, if $$\| \texttt{dstPoints} _i - \texttt{convertPointsHomogeneous} ( \texttt{H} \cdot \texttt{srcPoints} _i) \|_2 > \texttt{ransacReprojThreshold}$$ then the point $i$ is considered as an outlier. If srcPoints and dstPoints are measured in pixels, it usually makes sense to set this parameter somewhere in the range of 1 to 10. Preset to the OpenCV default (3.0).
maskoptNPARRAY,IMAGE,MASKOptional output mask set by a robust method ( RANSAC or LMeDS ). Note that the input mask values are ignored. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
maxItersoptINT2000-2147483648–2147483647The maximum number of RANSAC iterations. Preset to the OpenCV default (2000).
confidenceoptFLOAT0.9950-1e+38–1e+38Confidence level, between 0 and 1. The function finds and returns the perspective transformation $H$ between the source and the destination planes: $$s_i \vecthree{x'_i}{y'_i}{1} \sim H \vecthree{x_i}{y_i}{1}$$ so that the back-projection error $$\sum _i \left ( x'i- \frac{h{11} x_i + h_{12} y_i + h_{13}}{h_{31} x_i + h_{32} y_i + h_{33}} \right )^2+ \left ( y'i- \frac{h{21} x_i + h_{22} y_i + h_{23}}{h_{31} x_i + h_{32} y_i + h_{33}} \right )^2$$ is minimized. If the parameter method is set to the default value 0, the function uses all the point pairs to compute an initial homography estimate with a simple least-squares scheme. However, if not all of the point pairs ( $srcPoints_i$, $dstPoints_i$ ) fit the rigid perspective transformation (that is, there are some outliers), this initial estimate will be poor. In this case, you can use one of the three robust methods. The methods RANSAC, LMeDS and RHO try many different random subsets of the corresponding point pairs (of four pairs each, collinear pairs are discarded), estimate the homography matrix using this subset and a simple least-squares algorithm, and then compute the quality/goodness of the computed homography (which is the number of inliers for RANSAC or the least median re-projection error for LMeDS). The best subset is then used to produce the initial estimate of the homography matrix and the mask of inliers/outliers. Regardless of the method, robust or not, the computed homography matrix is refined further (using inliers only in case of a robust method) with the Levenberg-Marquardt method to reduce the re-projection error even more. The methods RANSAC and RHO can handle practically any ratio of outliers but need a threshold to distinguish inliers from outliers. The method LMeDS does not need any threshold but it works correctly only when there are more than 50% of inliers. Finally, if there are no outliers and the noise is rather small, use the default method (method=0). The function is used to find initial intrinsic and extrinsic matrices. Homography matrix is determined up to a scale. If $h_{33}$ is non-zero, the matrix is normalized so that $h_{33}=1$. Preset to the OpenCV default (0.995).

Outputs (2)

NameTypeDescription
HNPARRAY—
maskNPARRAY—