Nodes/ComfyUI CV/CV Rapid Track (Sequence)
ComfyUI Node

CV Rapid Track (Sequence)

An edge tracker, not a detector — and that's the whole catch

By bmad4ever·Created 4 months ago·Updated 15 days ago· 1
CV Rapid Track (Sequence)
  • frames
  • pts3d
  • tris
  • K
  • rvec
  • tvec
  • rvecs
  • tvecs
  • ratios
  • rmsds
  • rvec
  • tvec
  • tracked
◄num_control_points128►
◄search_length12►
◄iterations2►
◄scale1.00►

cv2.rapid implements the 1990 Harris & Stennett tracker, and it does something most modern tracking doesn't: it tracks a known 3D mesh through a video by matching the model's projected silhouette against image edges. You hand it geometry and an initial pose, and it hands you back a pose per frame. It's the machinery behind markerless AR, and it's in this pack because cv2.rapid is a class - the ~470 auto-generated cv2.* wrappers can't reach it.

The thing to understand before you wire it

It is a tracker, not a detector. Every frame it samples control points along the projected contour, searches perpendicular to the contour for the strongest image gradient, and refines the pose by PnP - starting from the previous frame's answer. That means a greedy loop with no global search:

  • Your initial pose has to already be close - tens of pixels, not "roughly in the frame". Get it from a marker, a solvePnP result, or the placement you built the scene with.
  • Motion between frames has to be small. Fast whip pans and motion blur will lose it.

If you need a detector, put an ArUco marker in the shot, solve that, and feed the answer here as the starting pose.

Inputs that matter

frames is an IMAGE batch and the batch is the sequence - frames are processed in order and the pose carries over, so don't shuffle or subsample it. pts3d + tris is your mesh: vertices and triangle indices. K is the camera matrix of the frames as given, and lens distortion is not modelled by cv2.rapid, so undistort wide-angle footage first. rvec / tvec are the initial pose for the first frame, in the mesh's own units.

Then the tuning, which is where this node gets opinionated and where the author's tooltips are worth reading rather than skimming:

  • num_control_points (128) - points sampled along the silhouette per iteration. Below ~32, a few bad correspondences can swing the pose.
  • search_length (12) - the capture range, in scaled pixels. It's not free: cv2.rapid takes the strongest gradient anywhere on the line, so a long search line lets background clutter outvote the silhouette. The author measured a textured vehicle over clutter pinning the error at ~46 px when this was raised from 12 to 96, no matter how good the start was.
  • scale (1.0) - the knob the author actually recommends for more capture range. Because the search range becomes search_length / scale in original pixels, and downscaling with INTER_AREA suppresses the texture edges that mislead a long line. It costs accuracy though: on a clip that starts on the object and only needs precision, dropping to 0.5 took measured error from 2.8 px to 18 px.
  • iterations (2) - refinement passes per frame. Two is the sweet spot. One lags; many passes drift on a static frame, because control points are re-extracted each pass and it isn't a converging solver.

Outputs

rvecs and tvecs are Bx3x1 stacks - one pose per frame, in order - and they're the point of the node. CV Rasterize Mesh takes that stack directly, so a tracked clip becomes a per-frame silhouette mask and a metric depth map in one more node.

ratios is the fraction of search lines that produced a usable correspondence: a relative health signal, and a sudden drop means the model came unstuck. rmsds is cv2's own 2D reprojection difference - read it together with ratios, because a low rmsd with a wrong pose is entirely possible when the tracker fits a self-consistent but shifted silhouette. tracked goes false if cv2.rapid failed on any frame; the poses stay valid up to that frame and repeat afterwards rather than halting your graph, so gate downstream compositing on it.

Install

Ships in ComfyUI CV (bmad4ever/comfyui_cv):

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"
# restart ComfyUI

Manager users: search the pack title. Needs Python ≥ 3.12 and a V3-API ComfyUI.

Where it breaks

Mesh density is a hidden input. cv2.rapid builds each 3D control point by interpolating linearly between two consecutive silhouette vertices, which is wrong between far-apart vertices - the pack measured up to 23 px of error on a raw truck mesh at fx 900. Its own docs are blunt: a coarse mesh makes the control points wrong by tens of pixels and the track diverges. Run the mesh through CV Mesh Split Long Edges first. The example workflows/42_rapid_model_tracking.json ships with a lane that bypasses that node specifically so you can watch it fail.

Winding matters. tris is used for backface culling and silhouette extraction, so an inside-out mesh tracks a ghost.

Set a frame budget. This is iterative CPU work over every frame; the batch is the sequence, so a 200-frame clip is 200 pose solves in one execution.

Categoryimage/CV/contrib

Inputs (10)

NameTypeDefaultDescription
framesIMAGEThe clip to track through, as one IMAGE batch (a video loader, or core 'Batch Images'). Frames are processed in order and the pose carries over, so the batch must BE the sequence.
pts3dNPARRAYNx3 model vertices, in the same units as tvec.
trisNPARRAYMx3 int triangle indices. Used for backface culling and the silhouette, so the winding matters.
KNPARRAY3x3 camera matrix of the FRAMES as given ('CV Camera Matrix' or a calibration). Lens distortion is not modelled by cv2.rapid - undistort the frames first if the lens is wide.
rvecNPARRAYInitial rotation (Rodrigues 3x1) for the FIRST frame. Typically from solvePnP on a marker, or the pose the object was placed at.
tvecNPARRAYInitial translation (3x1) for the first frame, in the mesh's units.
num_control_pointsINT1288–2048How many points to sample along the silhouette per iteration. More points average out bad matches but cost linearly. 128 is a good default; below ~32 a few wrong correspondences can swing the pose.
search_lengthINT121–256Half-length of the search line, in pixels of the SCALED frame - the tracker looks this far either side of the predicted contour. It sets the capture range, but NOT for free: cv2.rapid takes the strongest gradient anywhere on the line, so a long line lets texture and background edges outvote the silhouette. Measured on a textured vehicle over clutter, raising it from 12 to 96 pinned the error at ~46 px no matter how good the start was. Prefer raising capture range with 'scale' instead.
iterationsINT21–20Refinement passes per frame. 2 is usually the sweet spot; 1 lags behind the motion, and many passes on a static frame drift rather than settle (cv2.rapid's control points are re-extracted each pass, so it is not a converging solver).
scaleoptFLOAT1.000.05–1Downscale the frames before tracking (the camera matrix is scaled to match, so the pose stays in the original units). This is the CAPTURE-RANGE knob, and the one to reach for before search_length: the range it buys is search_length/scale in original pixels, it costs LESS than the equivalent search_length (the mask rasterization shrinks quadratically), and the INTER_AREA decimation suppresses the texture edges that mislead a long search line rather than adding more of them. It is NOT free accuracy: on a clip that starts on the object and only needs precision, dropping to 0.5 took the measured error from 2.8 px to 18 px. Leave it at 1.0 unless the pose you start from, or the motion between frames, is genuinely outside search_length.

Outputs (7)

NameTypeDescription
rvecsNPARRAYBx3x1 float64 stack: the rotation for each frame of the batch, in order. Slice one out with 'CV Index Batch' to project or render at that frame's pose.
tvecsNPARRAYBx3x1 float64 stack of translations, aligned with rvecs.
ratiosNPARRAYBx1 float32: per frame, the fraction of search lines that produced a usable correspondence (cv2.rapid's return value). A RELATIVE health signal - a sudden drop means the model came unstuck. Not an accuracy measure.
rmsdsNPARRAYBx1 float32: per frame, cv2.rapid's own 2D reprojection difference. Low rmsd with a wrong pose is possible (the tracker can fit a self-consistent but shifted silhouette), so read it together with ratios rather than alone.
rvecNPARRAYRotation after the LAST tracked frame - the starting pose for a following clip.
tvecNPARRAYTranslation after the last tracked frame.
trackedBOOLEANFalse if cv2.rapid failed on any frame (the model left the view, or its silhouette produced no usable vertices). The pose outputs are still valid up to that frame and repeat afterwards. Gate downstream compositing on it with an 'If/Else Switch'.