cv2.warpPerspective
The homography node, and what it's actually good at
- src
- M
- dsize
- result
A homography is the transform that maps four points to four points. That's it - and it's the reason this node exists: it's how you take a photo of a poster shot at an angle and get the poster back, flat and rectangular. It's also how you paste a patch into a plane in another image, un-keystone a picture of a screen, and rectify a stereo pair so the epipolar lines go horizontal.
cv2.warpAffine is the same node with a smaller matrix (2×3, parallel lines stay parallel). warpPerspective takes a 3×3 and lets parallel lines converge, which is the whole point - perspective is exactly the thing an affine warp can't do.
How it works
cv2 inverse-maps every destination pixel through the matrix and samples the source with your chosen interpolation. Because the destination geometry is what you specify, dsize decides the output canvas: for warpPerspective a dsize of (0, 0) in this wrapper means "same size as the source", which is fine when you're nudging an existing frame and wrong the moment you're rectifying a quad that doesn't fit the old aspect ratio.
Where the matrix comes from is the interesting part, and the pack ships all the usual sources: cv2.getPerspectiveTransform from four point pairs, cv2.findHomography (RANSAC over many correspondences, with the curated CV Find Homography (RANSAC) node above it), cv2.estimateAffine2D if affine is enough, or Parse Matrix if you're typing numbers from a paper.
Inputs that matter
src-IMAGE,MASK,NPARRAY, orLATENT. The socket is a match-type: whatever you feed in comes back as the same kind. AnIMAGEbatch is processed on frame 0 - this op isn't in the per-frame batch set, so unwind and re-stack (CV Unstack Batch/CV Stack Batch) if you need all frames.M- the 3×3 matrix. NPARRAY only; data in, not images.dsize- the output size as oneCV_TUPLE(w, h)value. For rectification this is the size of the rectangle you're warping to.flags- interpolation (INTER_LINEAR,INTER_CUBIC,INTER_LANCZOS4…) plus optionalWARP_INVERSE_MAP/WARP_FILL_OUTLIERS. The tooltip spells out the trap:WARP_INVERSE_MAP"sets M as the inverse transformation (dst → src)" - meaning cv2 treats your matrix as the forward map instead of the inverse it assumes by default.borderMode/borderValue- fill for the pixels that come from outside the source.BORDER_CONSTANTwith"(255, 255, 255)"gives you a white mat around a rectified document; leaveborderValueblank for OpenCV's zero.hint-ALGO_HINT_APPROXallows FP16 linear maths for speed where available.
Output is one value that echoes src's type, named result.
Latents, briefly
warpPerspective is on the pack's latent-safe list, so a LATENT link is warped in latent space - frame 0 becomes a float32 [H, W, C] array with the values untouched, no VAE round-trip and no 8-bit quantisation. That's more useful than it sounds when you want a slight perspective cue inside a sampler chain. It's also the only sane route for most latents, since cv2.rotate/cv2.transpose cap at 4 channels.
When the curated nodes are the better answer
This pack is one of the few places where the curated nodes are the ones to reach for first. CV Quad Warp takes four corners and does the rectification including the output size (which is the fiddly part). CV Paste Through Warp warps and composites onto a background in one step, premultiplied. CV Homography Map turns a homography into remap maps so you can compose it with other map generators. CV Affine Shape Warp does the rigid/affine case from correspondences. Use the raw wrapper when you want the plain function with your own matrix - a script, an experiment, a matrix produced by something upstream.
Installing it
Same pack, same dependency. Manager → search ComfyUI CV, or:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"
# restart ComfyUI
Python ≥3.12, recent V3-API ComfyUI.
What goes wrong
- A black frame. The most common cause is a matrix in the wrong direction - the warp maps everything outside the source. The second most common is
dsizetoo small to contain the result. - A stretched, wrong-proportion result. You got the matrix right and the output size wrong. Rectified output size is the diagonal-ish extent of the quad, not the source size; feed
CV Find Quadrilateral→CV Quad To Rectangleand let it tell you. - Integer matrices. cv2 wants floats. If you assembled
Mby hand as ints,CV Cast Arrayit first, or you'll get a runtime error that names nothing useful. (-215)assertions from cv2 surface through this pack as aRuntimeErrorthat names the function and the input shapes it received. Read that message before the traceback.
Inputs (7)
| Name | Type | Default | Description |
|---|---|---|---|
| src | COMFY_MATCHTYPE_V3 | input image. The image output(s) echo this input's format. A LATENT link is processed in latent space: frame 0 becomes a float32 [H,W,C] array (any channel count), values untouched. Arithmetic ops (add, multiply, etc.) also accept a full LATENT batch ({samples: [B,C,H,W]}) — the whole batch flows through when both inputs have the same batch size. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| M | NPARRAY | $3\times 3$ transformation matrix. A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here. | |
| dsize | CV_TUPLE | 0,0 | size of the output image. One value with 2 components (w, h) - it travels as a whole, so it cannot arrive half-connected. Wire it from 'CV Tuple' or type the components in place. |
| flagsopt | STRING | INTER_LINEAR | combination of interpolation methods (#INTER_LINEAR or #INTER_NEAREST) and the optional flag #WARP_INVERSE_MAP, that sets M as the inverse transformation ( $\texttt{dst}\rightarrow\texttt{src}$ ). cv2.warpPerspective flags: one of INTER_LINEAR, INTER_NEAREST, INTER_CUBIC, INTER_LANCZOS4 plus any of WARP_INVERSE_MAP, WARP_FILL_OUTLIERS, pipe-joined (e.g. "INTER_LINEAR | WARP_INVERSE_MAP"). In the UI this renders as a dropdown with one toggle per flag. |
| borderModeopt | COMBO | BORDER_DEFAULT | pixel extrapolation method (#BORDER_CONSTANT or #BORDER_REPLICATE). |
| borderValueopt | STRING | value used in case of a constant border; by default, it equals 0. cv2 Scalar as a literal, e.g. "(0, 255, 0)" (BGR) or "(0, 255, 0, 64)" (BGRA). A bare number broadcasts to every component, so "255" means (255, 255, 255, 255). Components past the target's channel count are ignored by OpenCV. Leave blank for the OpenCV default. | |
| hintopt | COMBO | ALGO_HINT_DEFAULT | Implementation modification flags. Set #ALGO_HINT_APPROX to use FP16 precision (if available) for linear calculation for faster speed. See #AlgorithmHint. |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| result | COMFY_MATCHTYPE_V3 | Echoes the 'src' input's format: an IMAGE link comes back as IMAGE, MASK as MASK, NPARRAY stays NPARRAY. |