Nodes/ComfyUI CV/CV Visual Odometry (Sequence)
ComfyUI Node

CV Visual Odometry (Sequence)

Recover a camera path from a video, scale problem included

By bmad4ever·Created 4 months ago·Updated 16 days ago· 1
CV Visual Odometry (Sequence)
  • frames
  • camera_matrix
  • scale_reference
  • trajectory
  • rotations
  • found_mask
  • inlier_counts
  • track_counts
  • redetections
  • found
◄max_features2000►
◄quality0.010►
◄min_distance7.00►
◄redetect_below1000►
◄win_size21►
◄max_level3►
◄fb_threshold1.0►
◄methodRANSAC►
◄threshold1.0►
◄confidence0.999►
◄min_inliers15►
◄scale_sourceconstant►
◄scale1.00►
◄motion_gatenone►
◄min_step0.00►

Point a phone out of a moving car window and there's a camera trajectory buried in the footage. CV Visual Odometry (Sequence) digs it out: it takes a whole video as an IMAGE batch and returns the camera centres, one row per frame, dead-reckoned from tracked features alone.

No depth model, no neural network, no training. Just the classical pipeline - corners, optical flow, essential matrix, composition - folded over an entire clip in a single node execution.

Why it's a sequence node and not a loop

This is the interesting design decision. You can build the same pipeline from single-step nodes: CV Track Features (KLT) → OpenCV Recover Pose (Essential Matrix) → CV Compose Pose (Trajectory), one instance per frame pair. That's a great way to see how odometry works on two frames.

It's a bad way to run a video, for two mechanical reasons. Frame-to-frame tracking wants the tracks from the previous pair to survive into the next one, and re-detecting corners every frame throws away the long tracks that make the pose stable. And a graph-level loop keeps one accumulator per iteration, while a pose update needs two (rotation and trajectory). So the fold doesn't close. This node does the loop in Python, where carrying two values between iterations is trivial, and calls the very same code the three single-step nodes use.

Between steps it keeps the surviving features, and only re-detects corners when their count drops below redetect_below (default 1000). Every step is the same trio: track, recover pose with RANSAC on the essential matrix, compose into the running pose.

The scale problem, up front

Monocular geometry fixes the direction of each step and never its length. There is no way around this - a single camera cannot tell a small nearby object from a large distant one, so the whole trajectory is only correct up to one global factor.

scale_source is how you resolve it:

  • constant (default) - every step is scale long. Shape right, size arbitrary.
  • per-step lengths (array) - one length per step, fed through scale_reference (wheel odometry, a known per-frame distance).
  • reference trajectory (step lengths) - scale_reference is an Nx3 track and its consecutive distances set the step lengths. This is how a monocular result becomes metric.

Inputs and outputs

Beyond the tracking parameters shared with CV Track Features (KLT) (max_features, quality, min_distance, win_size, max_level, fb_threshold), two gates are worth knowing. motion_gate with forward-dominant rejects steps whose translation is mostly sideways - the right guard for a forward-facing vehicle camera, and the wrong one for a drone that really does move sideways. min_step rejects very short steps: at a standstill the translation direction is pure noise, and integrating noise is drift with extra steps. Around 0.1 m at 10 fps for a car.

camera_matrix (the 3x3 K) is required and matters more than people expect. Without it there's no essential matrix, and an uncalibrated guess bends the whole path.

Outputs: trajectory (Nx3 camera centres, one row per input frame, row 0 at the origin - exactly what CV Trajectory Error (ATE / RPE) wants), rotations (Nx3x3 camera-to-world), found_mask (N uint8, 1 where the pose was really recovered and 0 where it was held), inlier_counts, track_counts (the health signal - a collapse here precedes a collapse in the pose), redetections, and found.

Failure tolerance is the design: a pair with no recoverable motion holds the previous pose and marks that frame 0 in found_mask. The trajectory always has exactly one row per input frame, which is what keeps everything downstream simple. Feed found_mask into CV Draw Path (ordered) as gap_mask so the guessed stretches are visible instead of silently drawn as real motion.

Installing it

Manager → ComfyUI CV, or:

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv

Restart. Python ≥ 3.12, a recent ComfyUI on the V3 node API, opencv-contrib-python-headless~=5.0.0.93. No models. GPL-3.0, forked from opencv-comfyui; the pack's own README says the example pipelines are tuned to specific datasets and not production-grade, and the visual-odometry workflow is explicitly one of them - it ships with a note that its 48 MB driving clip isn't in the repo (the ground-truth poses are, under a non-commercial licence).

Where people get burned

Expecting a usable scale from a single camera. You won't get one. This is the node's headline caveat, not a limitation to work around.

Rotation-only motion. Pan the camera from a fixed spot and there's no baseline to triangulate from - the pipeline has nothing to recover and holds the pose. Look at inlier_counts; low inliers across a whole clip usually means the footage was shot from one spot.

Textureless video. Blank sky, a wall, a tunnel. track_counts collapses, redetect_below fires constantly, the pose degrades. You'll see it in the numbers before you see it in the trajectory.

Long clips. Drift accumulates and there's no loop closure here - it's odometry, not SLAM. Over a few hundred frames a good run drifts a little; over a few thousand it drifts a lot, and no amount of parameter tuning fixes a problem that's structural.

Categoryimage/CV/features

Inputs (18)

NameTypeDefaultDescription
framesNPARRAY,IMAGEThe video as an IMAGE batch, in order ('Load Video' -> 'Get Video Components'). Fewer than 2 frames is a valid input: the result is a single origin and found=false.
camera_matrixNPARRAY3x3 intrinsic matrix K of the camera that shot the sequence ('CV Camera Matrix' or 'Calibrate Camera'). An uncalibrated guess bends the whole path - the essential matrix cannot be recovered without it.
max_featuresINT20000–20000Maximum corners to detect when a re-detection happens (goodFeaturesToTrack); 0 = every corner found. Higher means a slower but steadier pose.
qualityFLOAT0.0100.0001–1Minimum corner quality relative to the strongest corner (qualityLevel); lower keeps weaker corners.
min_distanceFLOAT7.001–100Minimum spacing in pixels between detected corners.
redetect_belowINT10005–20000Re-run corner detection once the surviving tracks drop under this count. Tracks die as the camera advances, so this is what keeps the pose fed; too low and the pose degrades before the refill, too high and it re-detects every frame (slower, and it throws away long tracks).
win_sizeINT213–101Lucas-Kanade search window (forced odd); larger tolerates bigger motion but blurs fine detail.
max_levelINT30–8Pyramid levels (0 = no pyramid); more levels track larger displacements - a fast-moving camera needs them.
fb_thresholdFLOAT1.00–30Forward-backward consistency gate in pixels: a track is dropped if re-tracking it back misses its origin by more than this. 0 disables the check.
methodCOMBORANSACRobust estimator for the essential matrix. RANSAC is the standard choice for tracked corners.
thresholdFLOAT1.00.1–100Max distance in pixels from a point to its epipolar line to count as an inlier (RANSAC).
confidenceFLOAT0.9990.5–1Desired probability that the estimate is correct.
min_inliersINT155–10000Minimum cheirality-consistent inliers for a step to count. Below it the step is rejected and the pose held.
scale_sourceCOMBOconstantWhere each step's LENGTH comes from, since the video cannot supply it. 'constant': every step is 'scale' long - the path's shape is right, its size is arbitrary. 'per-step lengths (array)': one length per step in 'scale_reference'. 'reference trajectory (step lengths)': 'scale_reference' is an Nx3 track (GPS, wheel odometry, ground truth) and its consecutive distances are used - this is how a monocular result is made metric.
scaleFLOAT1.000–1000000Step length for scale_source = 'constant'.
scale_referenceoptNPARRAYThe step lengths, as either an (N-1,) / (N,) array of lengths or an Nx3 reference trajectory. Shorter than the batch: the last value repeats. Ignored when scale_source = 'constant'.
motion_gateoptCOMBOnoneOptional sanity check on each recovered direction. 'forward-dominant' rejects a step whose translation is mostly sideways or vertical rather than along the optical axis - the classic guard for a forward-facing vehicle camera, where such a step is nearly always a bad essential-matrix fit. Leave at 'none' for a camera that really can move sideways (a drone, a handheld orbit).
min_stepoptFLOAT0.000–1000000Reject steps shorter than this (same unit as 'scale'). At a standstill the translation direction is pure noise, so integrating it only adds drift; 0.1 m is the usual choice for a car at 10 fps. 0 accepts every step.

Outputs (7)

NameTypeDescription
trajectoryNPARRAYCamera centres, Nx3 float64 - one row per input frame, row 0 at the origin. Feed 'CV Project To Plane'.
rotationsNPARRAYRunning camera-to-world rotations, Nx3x3 float64 (rotations[0] is the identity).
found_maskNPARRAY(N,) uint8 - 1 where that frame's pose came from a real recovered step, 0 where it was HELD because the step failed. Index 0 is 1 (the origin is known by definition). Feed it to 'CV Draw Path (ordered)' as 'gap_mask' so the held stretches are visible.
inlier_countsNPARRAY(N,) int32 cheirality inliers per step (0 at index 0).
track_countsNPARRAY(N,) int32 tracks that survived into each frame - the health signal to watch; a collapse here precedes a collapse in the pose.
redetectionsINTHow many times corners had to be re-detected.
foundBOOLEANTrue if at least one step was recovered. Gate the downstream preview on it with 'if/else'.