CV Compose Pose (Trajectory)
Chaining camera poses into a trajectory, with one number monocular vision can't give you
- rotation
- trajectory
- relative_rotation
- relative_translation
- rotation
- trajectory
- position
What this is for
Visual odometry: recover where the camera went, from video alone. The hard parts - matching features between frames, recovering the relative rotation and translation direction - are other nodes in this pack. This is the accumulation step, the boring one that turns a stack of relative poses into a path you can draw.
It's a loop-carried node. One instance per frame pair, each instance's rotation and trajectory outputs feeding the next one's inputs. If you've built any kind of iterative graph in ComfyUI you'll recognise the shape; if not, that's the whole structure - no hidden state, just the previous result wired back in.
How it works
CV Recover Pose hands you a relative rotation R and a unit translation t between two camera frames. The node updates the running camera-to-world rotation as R_wc ← R_wc · Rᵀ and appends the new camera centre as centre ← centre − scale · R_wc Rᵀ t, matching the convention that Recover Pose returns the mapping from the earlier frame to the later one.
The scale input is the honest bit, and it's worth understanding rather than just accepting: monocular geometry cannot recover absolute distance. A tiny object far away and a large object close by produce identical image motion, so the translation you get back is a direction only, normalised. The step length has to come from somewhere else - wheel odometry, a known baseline, a stereo pair, or (most often in practice) an assumption that the camera moved the same amount at every step. That's what scale is. Get it wrong and your trajectory is the right shape at the wrong size, which is exactly what it looks like: a perfect little map of your drive, 40% too long.
A trajectory that's empty or unconnected starts at the origin, so the first step in the chain doesn't need special handling.
Inputs and outputs
Required, and all five are:
rotation- the running 3×3 camera-to-world rotation. Start from identity:CV Camera Matrixwithfx=fy=1,cx=cy=0gives you one.trajectory- camera centres so far,N×3, last row = current centre. Start from a single origin point (CV Points).relative_rotation- the 3×3 fromCV Recover Pose.relative_translation- the 3×1 unit translation from the same node.scale(default 1.0) - absolute step length. See above; don't leave it at 1 and then wonder why the trajectory is measured in nothing.
Three outputs, two of which you feed back: rotation (updated 3×3) and trajectory ((N+1)×3) go to the next instance, and position is just the new camera centre as a 3×1 if you want to chart or threshold it without slicing the trajectory.
Wiring it up
Gate the chain on found from the pose-recovery step with an if/else, so a frame where the pose failed keeps the previous trajectory instead of appending garbage. Then project the finished trajectory with CV Project To Plane and draw it with CV Draw Polygon. Feeding the trajectory into CV Chart Scatter with y_down off is a quick way to eyeball a loop closure - if the path doesn't come back to where it started, something upstream is drifting and the length of that gap is your error.
Fair warning, straight from the pack's README: the visual-odometry example workflow is a demonstration, not a product. Its heuristics are fitted to specific datasets, and the driving clip it wants (48 MB) doesn't ship with the repo for size reasons - only the ground-truth poses do, under a non-commercial licence. Treat the example as a reference implementation of the pieces, not as something you point at your own footage and trust.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
# restart ComfyUI
Or ComfyUI CV in ComfyUI Manager (bmad4ever). Needs Python ≥ 3.12, a V3-node-API ComfyUI, and opencv-contrib-python-headless~=5.0.0.93. No models - this node is matrix arithmetic.
Common issues
- The trajectory is squashed flat. You're only chaining two frames, or
scaleis 1 while the real step is tiny and everything collapses into a dot. Set a plausible step length. - The path jumps wildly on some frames. Feature matching failed and returned a bad pose. Use the
foundflag upstream; the node itself has no way to know the pose it was handed is nonsense. - Nothing runs at all. The workflow needs another custom pack for its gating nodes -
ComfyUI-basic_data_handlingfor the IfElse node the pack's example uses. - Contrib nodes from the pack have vanished. A non-contrib OpenCV wheel emptied the contrib submodules;
python tools/repair_opencv_contrib.py --check, then--apply.
Inputs (5)
| Name | Type | Default | Description |
|---|---|---|---|
| rotation | NPARRAY | Running 3x3 camera-to-world rotation (identity for the first step). | |
| trajectory | NPARRAY | Camera centres so far, Nx3 (last row = current centre); start from a single origin point. | |
| relative_rotation | NPARRAY | 3x3 relative rotation R from 'CV Recover Pose'. | |
| relative_translation | NPARRAY | 3x1 UNIT relative translation t from 'CV Recover Pose'. | |
| scale | FLOAT | 1.000–1000000 | Absolute length of this step (monocular VO cannot recover it; use a known baseline or assume constant). |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| rotation | NPARRAY | Updated 3x3 camera-to-world rotation. |
| trajectory | NPARRAY | Camera centres with the new one appended, (N+1)x3. |
| position | NPARRAY | The new camera centre, 3x1. |