Nodes/ComfyUI CV/CV Annotate Correspondences
ComfyUI Node

CV Annotate Correspondences

Drag two images into alignment and get the warp for free

By bmad4ever·Created 3 months ago·Updated 14 days ago· 1
CV Annotate Correspondences
  • image
  • background
  • points_from
  • points_to
  • count
◄pairs[]►

Point correspondences are the input every geometry node wants and nobody enjoys producing. Automatic matching gives you hundreds of pairs and no idea which ones are wrong. This node gives you the manual version with a genuinely good editor bolted on: click a point where it is, drag the other end to where it should be, watch a live preview of the resulting warp, repeat until it looks right.

Why you'd reach for it

Because it's the authoring surface for every warp and fit in the pack, and because it removes the one class of bug that ruins all of them:

  • Feed a warp. The two outputs go straight into CV Thin Plate Spline Warp, CV Affine Shape Warp, CV Paste Through Warp, CV Find Homography (RANSAC), or the raw cv2.estimateAffine2D / cv2.getPerspectiveTransform wrappers.
  • Paste a subject into a scene. Connect the optional background and the whole framing flips: the TO end of every pair lives on the background, so the fit maps your image into the background's space. That's the authoring step for compositing, done by eye.
  • Author a deformation you can't describe. A face that should smile slightly, cloth that should drape differently - you click where things are and where they belong.

The bug class it removes is row misalignment. Each pair is one entry in one widget, so points_from and points_to are row-aligned by construction and cannot drift out of sync. Hand-built correspondences from two separate node outputs eventually will.

How it works

Run the workflow once and the image loads into the node's own canvas. Then:

  • LEFT-CLICK adds a pair. It starts pinned - from and to in the same place - and you drag either end.
  • Cyan is where a point IS (the from half); orange is where it must LAND (the to half); the arrow between them is the displacement, so you can see the motion field you're building at a glance.
  • SHIFT+CLICK removes a pair. ALT+DRAG moves both ends together (moving a whole correspondence without changing its displacement). DOUBLE-CLICK re-pins one end back onto the other if a drag went wrong.

With background connected, the canvas shows the two images side by side and the orange ends live on the background. Nothing about the fit changes - only which coordinate frame the to half is expressed in.

The pairs are stored in the pairs widget as normalized JSON, [[[from_x, from_y], [to_x, to_y]], ...], as fractions of the image size - so they survive a save, travel inside the workflow, and can be typed by hand when you need a point to land on an exact pixel.

The inputs and outputs

  • image - the source image, NPARRAY or IMAGE. Its size defines the coordinate system for the from half.
  • pairs - the JSON above. Maintained by dragging; editable by hand.
  • background (optional) - the destination image. When connected, the to half is in its pixels instead.
  • points_from - Nx1x2 float32 pixel coordinates in image. Always.
  • points_to - Nx1x2 float32 pixel coordinates: in background when one is connected, in image otherwise. Row-aligned with points_from.
  • count - how many pairs. Zero is valid - the node emits empty arrays rather than failing, so the graph runs before you've clicked.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv

Then restart ComfyUI and reload the page - the editor is front-end code, and a stale frontend is the usual reason the preview doesn't appear. Manager install works too (search ComfyUI CV). Dependency: opencv-contrib-python-headless~=5.0.0.93; Python 3.12+ and a ComfyUI on the V3 node API. It's an output node by design, so it executes in order to draw its preview.

Traps

  • The outputs are pixel coordinates of a specific size. Feed the consumer the same-size image. If you need to work at another resolution, warp first and resize after.
  • Consumer minimums still apply. Three pairs for an affine or a spline, four for a homography. With fewer, the downstream fit either errors or returns found = false and does nothing - the node will happily hand you two pairs.
  • Spread your pairs out. Five points clustered on one corner constrain a warp far less than five spread across the frame, and clusters are where a least-squares fit quietly goes wrong with a small residual.
  • The live preview is a preview. It's drawn client-side to be interactive; the authoritative result is whatever the fit node computes on the server. If they ever disagree, trust the server.
Categoryimage/CV/points

Inputs (3)

NameTypeDefaultDescription
imageNPARRAY,IMAGEThe image the correspondences are placed on; it is shown inside the node after the first run. The outputs are PIXEL coordinates of THIS image, so a consumer must be fed the same size. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
pairsSTRING[]Normalized [[[from_x, from_y], [to_x, to_y]], ...] JSON, 0-1 of the image size (of the background's size for the 'to' half when a background is connected). Maintained by dragging the preview below; hand-editing works as well.
backgroundoptNPARRAY,IMAGEOptional DESTINATION image. When connected, the editor shows it next to 'image', every pair's TO end lives on it, and points_to comes out in ITS pixel coordinates - feed points_from/points_to to a homography/affine fit to warp 'image' into place. Left unconnected, both halves live on 'image' as before. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.

Outputs (3)

NameTypeDescription
points_fromNPARRAYNx1x2 float32 PIXEL coordinates - where each control point currently is, in creation order. Always in 'image' pixels.
points_toNPARRAYNx1x2 float32 PIXEL coordinates - where the same-numbered control point must land. Row-aligned with points_from by construction. In 'background' pixels when one is connected, 'image' pixels otherwise.
countINT—