CV Track Features (KLT)
The corner tracker that tells you when it failed
- frame_a
- frame_b
- mask
- points_a
- points_b
- count
Give this node two frames of the same scene and it tells you which points moved where. That's it, and it's more than it sounds: tracking points between frames is the front half of monocular visual odometry, motion estimation, stabilisation, and anything else where you need the camera's motion rather than a diff image.
The node wraps the classic sparse tracker: find good corners in the first frame, chase them into the second with pyramidal Lucas-Kanade, then throw away the ones you don't believe.
How it works
Under the hood it's three OpenCV calls in a row. cv2.goodFeaturesToTrack finds up to max_features strong corners in frame_a. cv2.calcOpticalFlowPyrLK tracks each one into frame_b using a pyramid, so it handles a decent amount of displacement rather than only sub-pixel drift. Then a forward-backward check runs the tracks back into frame_a and drops any point that didn't land within fb_threshold pixels of where it started.
That last step is the good part, and it's why this node exists rather than you wiring the raw cv2_calcOpticalFlowPyrLK wrapper. Optical flow lies confidently - a corner that slid onto a different object or a repeating texture still comes back with status=1. The round-trip test is the cheapest honest way to catch it.
Both frames get grayscaled first, and frame_b is resized to frame_a if the sizes differ, so you don't have to match them yourself.
The inputs worth touching
- frame_a / frame_b - the earlier and later frame. Both accept a ComfyUI IMAGE or an NPARRAY. One caution: an IMAGE batch is treated as frame 0, so this node compares two single frames, not two videos.
- max_features (default 500) - how many corners to chase. More is slower and, in a textured scene, not necessarily better.
- quality (default 0.01) - the relative strength a corner needs to survive, as a fraction of the strongest one. Lower keeps weaker corners; raise it if you're getting junk.
- min_distance (default 7) - spacing between corners, in pixels. This is the anti-clumping knob; small values give you twenty corners all sitting on one doorknob.
- win_size (default 21) and max_level (default 3) - the search window and the pyramid depth. These are the two you raise when motion is fast.
- fb_threshold (default 1.0) - the consistency gate, in pixels.
0disables the check, which I wouldn't. - mask - optional, detect corners only where the mask is > 0. Handy to keep the tracker off the sky or off the moving car you don't want it following.
Outputs are points_a and points_b (both Nx1x2 float32, same order) plus count. Wire the pair straight into CV Recover Pose (Essential Matrix) for the camera motion, or into CV Triangulate Points (Two-View) if you want 3D out of it.
Installing it
Manager → search ComfyUI CV, or:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
Restart ComfyUI afterwards. The pack wants Python ≥ 3.12 and a current ComfyUI on the V3 node API, and pulls in opencv-contrib-python-headless~=5.0.0.93 plus numpy and torch. It's a fork of geroldmeisinger's opencv-comfyui, GPL-3.0, and the author is upfront that it's a personal project with LLM-written code and no support promise - one more reason to keep an eye on the outputs rather than trusting them.
No model files needed for this node.
Where people get burned
Count = 0 is not a bug. A blank frame, a textureless wall, heavy motion blur - you get empty arrays and count of 0. The node deliberately doesn't raise, so a visual-odometry pipeline doesn't die on one bad frame. Don't treat it as an error; branch on it.
Fast motion kills it silently. Pyramidal LK tolerates real displacement, but not unlimited. If your count collapses between two frames of a whip pan, raise max_level to 4-5 and bump win_size. There's a cost: a big window smooths detail, so you'll track a blurrier point.
Resolution matters more than you'd think. Tracking on a 256 px thumbnail gives you scale-free mush. Track full-res and scale the points afterwards if you need to.
A mask that's the wrong size. The mask has to line up with the frame; a mismatched mask is one of those failures that shows up as "zero corners and no explanation".
Two frames, not a sequence. For a whole video, the single-pair version is for understanding the pipeline - debugging on frames you pick. Use CV Visual Odometry (Sequence) to run the same code folded over an entire IMAGE batch, because it carries surviving tracks from frame to frame instead of re-detecting every time.
Inputs (9)
| Name | Type | Default | Description |
|---|---|---|---|
| frame_a | NPARRAY,IMAGE | Earlier frame; corners are detected here. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| frame_b | NPARRAY,IMAGE | Later frame; resized to frame_a if sizes differ. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| max_features | INT | 5000–10000 | Maximum corners to detect in frame_a (goodFeaturesToTrack); 0 = return all detected corners. |
| quality | FLOAT | 0.0100.0001–1 | Minimum corner quality relative to the strongest corner (qualityLevel); lower keeps weaker corners. |
| min_distance | FLOAT | 7.001–100 | Minimum spacing in pixels between detected corners. |
| win_size | INT | 213–101 | Lucas-Kanade search window (forced odd); larger tolerates bigger motion but blurs fine detail. |
| max_level | INT | 30–8 | Pyramid levels (0 = no pyramid); more levels track larger displacements. |
| fb_thresholdopt | FLOAT | 1.00–30 | Forward-backward consistency gate in pixels: a track is dropped if re-tracking it back misses its origin by more than this. 0 disables the check. |
| maskopt | NPARRAY,MASK | Optional mask: detect corners only where mask > 0. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| points_a | NPARRAY | Tracked corners in frame_a, Nx1x2 float32. |
| points_b | NPARRAY | Their tracked positions in frame_b, Nx1x2 float32 (same order). |
| count | INT | — |