Nodes/ComfyUI CV/cv2.perspectiveTransform
ComfyUI Node

cv2.perspectiveTransform

Push points through a homography instead of resampling an image

By bmad4ever·Created 4 months ago·Updated 14 days ago· 1
cv2.perspectiveTransform
  • src
  • m
  • nparray

warpPerspective moves pixels. This moves points - same matrix, no resampling, no image. Which is exactly what you want when you've already got a homography and you need to know where a corner, a landmark, a bounding box or a set of keypoints lands on the other side of it.

Two inputs, both NPARRAY-only (the wrapper refuses IMAGE/MASK links here on purpose - a picture is not a point set), one NPARRAY out.

  • src - points. "Input two-channel or three-channel floating-point array; each element is a 2D/3D vector to be transformed." In practice that's N x 1 x 2 from CV Points, CV Contour To Points, CV Detect Corners, CV ArUco Detect Markers (corners) or any feature detector, or N x 1 x 3 for 3D.
  • m - the transform. "3x3 or 4x4 floating-point transformation matrix." A 3x3 for 2D points, a 4x4 for 3D. CV Find Homography (RANSAC) emits the 3x3; CV Pose To Matrix and Parse Matrix are the other two places it comes from.

Output: nparray, points in the same layout you sent them.

What's actually happening

A projective transform on points is a matrix multiply followed by a divide: [x', y', w'] = M · [x, y, 1], then x'/w', y'/w'. That divide is the "perspective" part - it's what makes parallel lines converge, and it's why m's bottom row matters. It's also why points behind the camera come back with negative w and produce coordinates that look plausible but are on the wrong side of the plane. Nothing clips, nothing warns.

That last point is the difference from the image version. warpPerspective has a dsize and clamps its output to a canvas; perspectiveTransform has neither. Points that land off-image are just numbers - negative, out of range, whatever the math says. So it's on you to check, and the pack's CV Filter Points By Distance / CV Filter Points By Mask nodes are the tools for trimming what you don't want.

Where you'd use it

Following a box between image spaces. You've detected something in a resized or rectified frame and you want its box in original coordinates. If the mapping is scale + shift, CV Scale BBoxes does it in one node. If it's projective, take the box corners with CV BBoxes To Array, run them through this node, and convert the outline back with CV Array To BBoxes.

Stitching and template matching. CV Match Template Multi-Scale gives you a crop and a scale; the matrix that puts it back in the source's frame is a perspectiveTransform call away from being useful. Same for the panorama nodes: CV Stitch Canvas produces warp_matrix for exactly this kind of point bookkeeping.

Moving between camera views. With a homography from CV Find Homography (RANSAC) or a rigid pose from CV Pose To Matrix, this node answers "where does this point appear in the other camera" without rendering either image. If you're doing plane-based AR-style overlays or checking a triangulation, that's the whole job.

3D, with the 4x4 form. Feed N x 1 x 3 3D points and a 4x4 (a pose plus translation) and you get transformed 3D points back - a rigid transform without going through cv2.transform or a point-cloud node. The pack's CV Transform Points 3D is the dedicated version of that idea if you want a friendlier interface.

The errors you'll actually see

Wrong point layout. cv2 wants N x 1 x 2 (or N x 1 x 3). An N x 2 array of floats, or an int32 one, gives you the Python binding's catch-all "Overload resolution failed" with no hint about which argument it hates. CV Points and CV Contour To Points emit the right thing; hand-built arrays are where this goes wrong.

Mismatched matrix and point dimension. A 4x4 with N x 1 x 2 input is an error, not a 2D transform. Check with CV Array Shape if the matrix came from somewhere you don't fully trust.

Integers where floats belong. Everything here is float32/float64 - CV Cast Array first if any upstream node produced ints (contour points often are).

There's no batch behaviour to speak of: this is a matrix node. Link an image into it if you insist and it takes frame 0, which for a corner-detection input means you're transforming corners you found on one frame.

Installing

ComfyUI Manager → search ComfyUI CV → Install → restart, or:

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"

Python ≥ 3.12, a recent ComfyUI (the pack's wrappers are V3 node definitions, generated at import time from whatever functions your OpenCV exposes), and no models. perspectiveTransform is core cv2, so no contrib requirement - the node is under image/CV/low-level/cv2 P.

Categoryimage/CV/low-level/cv2 P

Inputs (2)

NameTypeDefaultDescription
srcNPARRAYinput two-channel or three-channel floating-point array; each element is a 2D/3D vector to be transformed. A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here.
mNPARRAY3x3 or 4x4 floating-point transformation matrix. A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here.

Outputs (1)

NameTypeDescription
nparrayNPARRAY—