CV Project Points (Sequence)
One point set, a whole tracked clip
- points_3d
- rvecs
- tvecs
- K
- dist_coeffs
- scene_depth
- points_2d
- visible
- depths
The "and now draw it" node of the 3D tracking pipeline. CV Rapid Track (Sequence) tracks a mesh through a clip and hands you a stack of rotations and translations; this projects a fixed set of 3D points through that whole stack, so a label or an anchor follows the object across every frame.
What it's for
Annotations and measurements that live on a moving object. You pick anchor points on a model - the tip of a nozzle, the four corners of a panel, two ends of a dimension you're measuring - and this node tells you where those points are in each frame. points_3d is Nx3 in the mesh's object space, so anchors stay glued to the geometry rather than to the image.
The alternative is a cv2.projectPoints call per frame, which in a graph means either a loop you have to manage or N near-identical nodes. This takes the Bx3x1 rvecs and tvecs as arrays and projects the whole batch in one go.
The depth test, which is the interesting half
Wire scene_depth - a BxHxW (or HxW) metric depth map from CV Rasterize Mesh - and every projected point is compared against the depth actually rendered at its pixel. A point on the far side of the object, behind the surface the camera can see, gets flagged. That's hidden-line removal, the difference between an annotation that looks attached to the object and one that looks pasted on top of it.
What happens to invisible points is hidden_points. Keep coordinates returns where the point would have been - right for measuring, or for your own gating. Blank out (NaN) replaces them, and because the drawing nodes skip non-finite points you get the hidden-line effect for free without touching your labelling: rows stay index-aligned with their captions, which filtering them out would break by renumbering everything. depth_tolerance (0.02, in scene units) says how far behind the rendered surface a point may sit and still count as visible - an anchor picked on the surface lands within about a rasterisation pixel, so a small tolerance absorbs that without letting genuine far-side points through.
Inputs and outputs
points_3d, rvecs, tvecs, K are required. A single 3x1 rotation works as well as a Bx3x1 stack. dist_coeffs is optional - leave it unwired for a pinhole camera, and note that CV Rasterize Mesh is pinhole-only, so a distorted projection won't line up with its depth map. Undistort the frames instead of mixing the two.
Out comes points_2d, a BxNx2 float32 array of pixel positions, one row per pose; slice a frame with CV Index Batch and feed CV Draw Points, CV Draw Labels or CV Draw Connections. visible is a BxN uint8 mask, all ones when scene_depth isn't wired (nothing to test against) and zero for points that fall outside the frame - gate the drawing on it with CV Filter Points By Mask. depths is the camera-space Z of every point in scene units, which is how you scale a label with distance or sort overlapping callouts back to front.
Installing
ComfyUI Manager → search "ComfyUI CV", or:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
# restart ComfyUI
Python ≥ 3.12, recent ComfyUI on the V3 node API, and the pinned contrib OpenCV:
pip install "opencv-contrib-python-headless~=5.0.0.93"
Gotchas
Where the anchors come from matters more than how they're projected. Picking points on a model with CV Pick Point On Mesh or CV Annotate Points On Model gives you coordinates in the model's own frame, which is what this node expects. Typing coordinates by eye of an arbitrary GLB almost never works, because you don't know where its origin is.
Depth and distortion don't mix. If the projected points don't sit on the object even though the pose is right, check whether you're comparing a lens-distorted projection against a pinhole depth map. Pick one.
Rasterise once. depth_metric from CV Preview 3D (Calibrated Camera) carries the lens distortion and is already registered with the render, so if you've rendered the model for an overlay you can reuse that z-buffer here instead of rasterising a second time.
Inputs (8)
| Name | Type | Default | Description |
|---|---|---|---|
| points_3d | NPARRAY | Nx3 points in the mesh's object space - annotation anchors, measurement ends, a bounding box's corners. Pick them off the model with the 'CV Pick Point On Mesh' blueprint rather than typing coordinates. | |
| rvecs | NPARRAY | Bx3x1 rotations (a single 3x1 works too). | |
| tvecs | NPARRAY | Bx3x1 translations, same length as rvecs. | |
| K | NPARRAY | 3x3 camera matrix of the frames. | |
| dist_coeffsopt | NPARRAY | Optional lens distortion (k1, k2, p1, p2[, k3...]). Leave unwired for a pinhole camera. Note that 'CV Rasterize Mesh' is pinhole-only, so a distorted projection will NOT line up with its depth map - undistort the frames instead of mixing the two. | |
| scene_depthopt | NPARRAY | Optional BxHxW (or HxW) metric depth from 'CV Rasterize Mesh'. When wired, each projected point is compared with the depth actually rendered at its pixel and reported in 'visible'. | |
| hidden_pointsopt | COMBO | keep coordinates | What points_2d holds where 'visible' is 0. 'keep coordinates' returns where the point WOULD be, which is what you want for measuring or for your own gating. 'blank out (NaN)' replaces them with NaN - the drawing nodes skip non-finite points, so an overlay gets hidden-line removal for free and the rows stay index-aligned with their labels (filtering them out instead renumbers everything). Only ever differs when scene_depth is wired. |
| depth_toleranceopt | FLOAT | 0.0200–1000 | How far BEHIND the rendered surface a point may sit and still count as visible, in scene units. An anchor picked ON the surface lands within a rasterization pixel of it, so a small tolerance (~1% of the object) absorbs that without letting far-side points through. |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| points_2d | NPARRAY | BxNx2 float32 pixel positions, one row per pose. Slice a frame out with 'CV Index Batch' and feed 'CV Draw Points' / 'Draw Labels' / 'Draw Connections'. |
| visible | NPARRAY | BxN uint8: 1 where the point faces the camera, by comparing its own Z with scene_depth at its pixel. All ones when scene_depth is not wired (nothing to test against) and 0 for points that fall outside the frame. Gate the drawing on it with 'CV Filter Points By Mask'. |
| depths | NPARRAY | BxN float32 camera-space Z of each point - how far it is from the camera, in scene units. Useful to scale a label with distance, or to sort overlapping callouts back to front. |