CV Visual Odometry (Sequence)
Recover a camera path from a video, scale problem included
- frames
- camera_matrix
- scale_reference
- trajectory
- rotations
- found_mask
- inlier_counts
- track_counts
- redetections
- found
Point a phone out of a moving car window and there's a camera trajectory buried in the footage. CV Visual Odometry (Sequence) digs it out: it takes a whole video as an IMAGE batch and returns the camera centres, one row per frame, dead-reckoned from tracked features alone.
No depth model, no neural network, no training. Just the classical pipeline - corners, optical flow, essential matrix, composition - folded over an entire clip in a single node execution.
Why it's a sequence node and not a loop
This is the interesting design decision. You can build the same pipeline from single-step nodes: CV Track Features (KLT) → OpenCV Recover Pose (Essential Matrix) → CV Compose Pose (Trajectory), one instance per frame pair. That's a great way to see how odometry works on two frames.
It's a bad way to run a video, for two mechanical reasons. Frame-to-frame tracking wants the tracks from the previous pair to survive into the next one, and re-detecting corners every frame throws away the long tracks that make the pose stable. And a graph-level loop keeps one accumulator per iteration, while a pose update needs two (rotation and trajectory). So the fold doesn't close. This node does the loop in Python, where carrying two values between iterations is trivial, and calls the very same code the three single-step nodes use.
Between steps it keeps the surviving features, and only re-detects corners when their count drops below redetect_below (default 1000). Every step is the same trio: track, recover pose with RANSAC on the essential matrix, compose into the running pose.
The scale problem, up front
Monocular geometry fixes the direction of each step and never its length. There is no way around this - a single camera cannot tell a small nearby object from a large distant one, so the whole trajectory is only correct up to one global factor.
scale_source is how you resolve it:
- constant (default) - every step is
scalelong. Shape right, size arbitrary. - per-step lengths (array) - one length per step, fed through
scale_reference(wheel odometry, a known per-frame distance). - reference trajectory (step lengths) -
scale_referenceis anNx3track and its consecutive distances set the step lengths. This is how a monocular result becomes metric.
Inputs and outputs
Beyond the tracking parameters shared with CV Track Features (KLT) (max_features, quality, min_distance, win_size, max_level, fb_threshold), two gates are worth knowing. motion_gate with forward-dominant rejects steps whose translation is mostly sideways - the right guard for a forward-facing vehicle camera, and the wrong one for a drone that really does move sideways. min_step rejects very short steps: at a standstill the translation direction is pure noise, and integrating noise is drift with extra steps. Around 0.1 m at 10 fps for a car.
camera_matrix (the 3x3 K) is required and matters more than people expect. Without it there's no essential matrix, and an uncalibrated guess bends the whole path.
Outputs: trajectory (Nx3 camera centres, one row per input frame, row 0 at the origin - exactly what CV Trajectory Error (ATE / RPE) wants), rotations (Nx3x3 camera-to-world), found_mask (N uint8, 1 where the pose was really recovered and 0 where it was held), inlier_counts, track_counts (the health signal - a collapse here precedes a collapse in the pose), redetections, and found.
Failure tolerance is the design: a pair with no recoverable motion holds the previous pose and marks that frame 0 in found_mask. The trajectory always has exactly one row per input frame, which is what keeps everything downstream simple. Feed found_mask into CV Draw Path (ordered) as gap_mask so the guessed stretches are visible instead of silently drawn as real motion.
Installing it
Manager → ComfyUI CV, or:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
Restart. Python ≥ 3.12, a recent ComfyUI on the V3 node API, opencv-contrib-python-headless~=5.0.0.93. No models. GPL-3.0, forked from opencv-comfyui; the pack's own README says the example pipelines are tuned to specific datasets and not production-grade, and the visual-odometry workflow is explicitly one of them - it ships with a note that its 48 MB driving clip isn't in the repo (the ground-truth poses are, under a non-commercial licence).
Where people get burned
Expecting a usable scale from a single camera. You won't get one. This is the node's headline caveat, not a limitation to work around.
Rotation-only motion. Pan the camera from a fixed spot and there's no baseline to triangulate from - the pipeline has nothing to recover and holds the pose. Look at inlier_counts; low inliers across a whole clip usually means the footage was shot from one spot.
Textureless video. Blank sky, a wall, a tunnel. track_counts collapses, redetect_below fires constantly, the pose degrades. You'll see it in the numbers before you see it in the trajectory.
Long clips. Drift accumulates and there's no loop closure here - it's odometry, not SLAM. Over a few hundred frames a good run drifts a little; over a few thousand it drifts a lot, and no amount of parameter tuning fixes a problem that's structural.
Inputs (18)
| Name | Type | Default | Description |
|---|---|---|---|
| frames | NPARRAY,IMAGE | The video as an IMAGE batch, in order ('Load Video' -> 'Get Video Components'). Fewer than 2 frames is a valid input: the result is a single origin and found=false. | |
| camera_matrix | NPARRAY | 3x3 intrinsic matrix K of the camera that shot the sequence ('CV Camera Matrix' or 'Calibrate Camera'). An uncalibrated guess bends the whole path - the essential matrix cannot be recovered without it. | |
| max_features | INT | 20000–20000 | Maximum corners to detect when a re-detection happens (goodFeaturesToTrack); 0 = every corner found. Higher means a slower but steadier pose. |
| quality | FLOAT | 0.0100.0001–1 | Minimum corner quality relative to the strongest corner (qualityLevel); lower keeps weaker corners. |
| min_distance | FLOAT | 7.001–100 | Minimum spacing in pixels between detected corners. |
| redetect_below | INT | 10005–20000 | Re-run corner detection once the surviving tracks drop under this count. Tracks die as the camera advances, so this is what keeps the pose fed; too low and the pose degrades before the refill, too high and it re-detects every frame (slower, and it throws away long tracks). |
| win_size | INT | 213–101 | Lucas-Kanade search window (forced odd); larger tolerates bigger motion but blurs fine detail. |
| max_level | INT | 30–8 | Pyramid levels (0 = no pyramid); more levels track larger displacements - a fast-moving camera needs them. |
| fb_threshold | FLOAT | 1.00–30 | Forward-backward consistency gate in pixels: a track is dropped if re-tracking it back misses its origin by more than this. 0 disables the check. |
| method | COMBO | RANSAC | Robust estimator for the essential matrix. RANSAC is the standard choice for tracked corners. |
| threshold | FLOAT | 1.00.1–100 | Max distance in pixels from a point to its epipolar line to count as an inlier (RANSAC). |
| confidence | FLOAT | 0.9990.5–1 | Desired probability that the estimate is correct. |
| min_inliers | INT | 155–10000 | Minimum cheirality-consistent inliers for a step to count. Below it the step is rejected and the pose held. |
| scale_source | COMBO | constant | Where each step's LENGTH comes from, since the video cannot supply it. 'constant': every step is 'scale' long - the path's shape is right, its size is arbitrary. 'per-step lengths (array)': one length per step in 'scale_reference'. 'reference trajectory (step lengths)': 'scale_reference' is an Nx3 track (GPS, wheel odometry, ground truth) and its consecutive distances are used - this is how a monocular result is made metric. |
| scale | FLOAT | 1.000–1000000 | Step length for scale_source = 'constant'. |
| scale_referenceopt | NPARRAY | The step lengths, as either an (N-1,) / (N,) array of lengths or an Nx3 reference trajectory. Shorter than the batch: the last value repeats. Ignored when scale_source = 'constant'. | |
| motion_gateopt | COMBO | none | Optional sanity check on each recovered direction. 'forward-dominant' rejects a step whose translation is mostly sideways or vertical rather than along the optical axis - the classic guard for a forward-facing vehicle camera, where such a step is nearly always a bad essential-matrix fit. Leave at 'none' for a camera that really can move sideways (a drone, a handheld orbit). |
| min_stepopt | FLOAT | 0.000–1000000 | Reject steps shorter than this (same unit as 'scale'). At a standstill the translation direction is pure noise, so integrating it only adds drift; 0.1 m is the usual choice for a car at 10 fps. 0 accepts every step. |
Outputs (7)
| Name | Type | Description |
|---|---|---|
| trajectory | NPARRAY | Camera centres, Nx3 float64 - one row per input frame, row 0 at the origin. Feed 'CV Project To Plane'. |
| rotations | NPARRAY | Running camera-to-world rotations, Nx3x3 float64 (rotations[0] is the identity). |
| found_mask | NPARRAY | (N,) uint8 - 1 where that frame's pose came from a real recovered step, 0 where it was HELD because the step failed. Index 0 is 1 (the origin is known by definition). Feed it to 'CV Draw Path (ordered)' as 'gap_mask' so the held stretches are visible. |
| inlier_counts | NPARRAY | (N,) int32 cheirality inliers per step (0 at index 0). |
| track_counts | NPARRAY | (N,) int32 tracks that survived into each frame - the health signal to watch; a collapse here precedes a collapse in the pose. |
| redetections | INT | How many times corners had to be re-detected. |
| found | BOOLEAN | True if at least one step was recovered. Gate the downstream preview on it with 'if/else'. |