Nodes/ComfyUI CV/CV Annotate Points
ComfyUI Node

CV Annotate Points

Click a seed, get a mask — the node most graphs secretly need

By bmad4ever·Created 3 months ago·Updated 14 days ago· 1
CV Annotate Points
  • image
  • points
  • count
  • mask
◄points[]►

Almost every mask-editing story in ComfyUI starts with "load the image, run it once, then click". This node is that click: it renders the image inside itself, you left-click to drop points, and you get three things out - the points in pixel coordinates, how many there are, and a filled polygon mask. It's the manual-input end of a pack that has a lot of automatic ends.

Why you'd reach for it

  • Four corners → CV Quad Warp. Click the four corners of a screen, a page, a poster, and the quad warp does the perspective. This is the fastest path from "photo of a whiteboard" to "straightened image of the whiteboard" in any node pack.
  • A seed point → CV Select Component At Point. Grabcut-ish and connected-component work is all "start from here": the workflow needs a click, not a prompt.
  • A polygon mask → anything. The mask output is just a filled polygon, so it can drive an inpaint region, a crop, a local detail pass, a colour-correction region.
  • An outline → CV Thin Plate Spline Warp. Places points, not correspondences; if you want a warp, compare with CV Annotate Correspondences, which pairs them.

How it works, and the editing gestures that are actually the point

Run the workflow once and the image appears inside the node. Then:

  • Left-click adds a point. Drag moves the nearest one.
  • SHIFT+CLICK removes the nearest point.
  • Once the outline is closed (3 or more points), a click is inserted into the nearest edge - highlighted so you can see which - instead of being appended at the end. That's the difference between refining a shape where it's wrong and re-clicking the whole thing, and it's the reason this node is pleasant rather than infuriating.
  • ALT+CLICK opts out of that and appends at the end, in click order, when order matters to the consumer.

Under the hood the points live in the points widget as normalized JSON, [[x, y], ...] as fractions of the image size - which means they survive a save/reload, travel with the workflow, and can be typed or pasted by hand when you want exact values rather than a click that landed two pixels off.

The inputs and outputs

Inputs are just two:

  • image - the canvas, and the coordinate system: NPARRAY or IMAGE. Batch frames beyond the first aren't what you're annotating, so treat it as a single image.
  • points - the JSON above. You normally never touch it; it's updated by the preview.

Outputs:

  • points - Nx1x2 float32 pixel coordinates, in click order.
  • count - how many. Zero is a valid, non-error state: the node emits an empty array rather than failing, so a graph can execute before you've clicked anything.
  • mask - a uint8 image-sized mask, 255 inside the polygon. Fewer than 3 points and it's all zeros.

Note that mask is an NPARRAY, not a ComfyUI MASK - convert it with CV Array → Mask before it meets a core mask node. That conversion step is a pack convention worth internalising: the CV nodes work on raw ndarrays and the graph works on typed tensors, and the pack keeps the two ends clearly separated.

Install

ComfyUI Manager → ComfyUI CV, or the usual:

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv

Restart ComfyUI, and reload the browser page - the annotation widget is front-end code, so a hard refresh after installing is not superstition. Requirements: opencv-contrib-python-headless~=5.0.0.93, Python 3.12+, and a ComfyUI new enough for the V3 node API.

Traps

  • Pixel coordinates are tied to a size. The outputs are in the pixel space of the image you fed in. If a downstream consumer is working on a differently-sized image, the points will be wrong in a way that looks almost right. Resize after the warp, not before the fit.
  • Points are stored normalized, so a resize is survivable - the JSON stays 0–1 and re-resolves against whatever the image now is. That's a feature, but it also means a workflow that swaps in a cropped image will move your points without telling you.
  • The node is an output node. It runs so it can draw its preview; that's intentional, not a misconfiguration, and it's why the preview shows up without you wiring anything to a save.
  • Hand-editing JSON is fiddly but worth it for a repeated task: type the four corners of a known template once, save the workflow, and every future run starts pre-clicked.
Categoryimage/CV/points

Inputs (2)

NameTypeDefaultDescription
imageNPARRAY,IMAGEThe image to annotate; it is shown inside the node after the first run. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
pointsSTRING[]Normalized [[x, y], ...] JSON, 0-1 of the image size. Maintained by clicking the preview below; hand-editing works as well.

Outputs (3)

NameTypeDescription
pointsNPARRAYNx1x2 float32 PIXEL coordinates, in click order.
countINT—
maskNPARRAYuint8 image-sized mask, 255 inside the polygon the points enclose (click order, any count >= 3; fewer points -> all zeros).