Nodes/ComfyUI CV/cv2.recoverPose (1/4)
ComfyUI Node

cv2.recoverPose (1/4)

Relative camera pose from two calibrated views

By bmad4ever·Created 3 months ago·Updated 14 days ago· 1
cv2.recoverPose (1/4)
  • points1
  • points2
  • cameraMatrix1
  • distCoeffs1
  • cameraMatrix2
  • distCoeffs2
  • mask
  • retval
  • E
  • R
  • t
  • mask
◄methodRANSAC►
◄prob0.9990►
◄threshold1.0000►

What it's for

Two photos of the same static scene, taken by two different cameras (or one camera before and after you moved it). You've matched features between them. This node answers: how did the camera move? - the rotation, the direction of translation, and a mask saying which matches were trusted.

That's the core of two-view structure-from-motion, and it's the second half of monocular visual odometry: match features, then recover pose per frame pair, then chain the poses into a trajectory. In this pack it sits underneath the curated CV Recover Pose (Essential Matrix) node and the CV Visual Odometry (Sequence) node - the (1/4) suffix tells you it's the first of four overloads, and this one is the most general: separate intrinsics and distortion for each camera.

How it works

The pose isn't read off the points directly. cv2 estimates an essential matrix from the correspondences and then decomposes it. That decomposition has four candidate solutions, and three of them are geometrically impossible because they put the scene behind the camera. The cheirality check throws those away: it keeps the rotation/translation for which the inliers triangulate to positive depth in both views. That's the node's real content - the essential-matrix plumbing is done for you here, unlike the other three overloads which take an E you estimated elsewhere.

The mask output is what you filter with: only inliers that survive the chirality check are marked, so a wire into CV Filter Points By Mask (or the pack's draw-matches node) is how you see whether the estimate is any good.

The inputs that matter

Six arrays are required - points1, points2 (matching Nx2 point sets, float coordinates, from CV Match Features), then cameraMatrix1, distCoeffs1, cameraMatrix2, distCoeffs2. Get K and distortion from a calibration (CV Calibrate Camera (ChArUco) / CV Camera Matrix) or from CV Load Camera Params (JSON). If both cameras are the same camera, feed the same K twice; that's the common case and it's fine.

Optional:

  • method - RANSAC by default, or LMEDS. RANSAC is the normal choice for feature matches.
  • prob - 0.999, the confidence RANSAC aims for.
  • threshold - 1.0 px, the distance from a point to its epipolar line beyond which it's an outlier. Bump it to 2–3 if your matches are sloppy or your images are large and noisy.
  • mask - an input mask of matches to consider. Feed the output mask back in if you're re-running on the same pair and want to lock onto the inliers you already trust.

Outputs: retval (count of inliers used, so gate on it), E (the essential matrix), R (3×3 rotation), t (3×1 translation), mask (N×1 uint8 inlier mask).

And the sentence that bites everyone: t is a unit vector. Monocular geometry fixes the direction of the camera's motion, never the distance. If you need metres you need a known baseline (stereo), a known object size, or an IMU. Everything downstream inherits that scale ambiguity - pose trajectories from this node come out in arbitrary units, which is why the pack ships CV Trajectory Error (ATE/RPE) to compare one against another after an Umeyama alignment.

Wiring it downstream

R and t go into CV Triangulate Points (Two-View) for sparse 3D, or into CV Compose Pose (Trajectory) to accumulate a path, or cv2.Rodrigues → rvec when something wants a rotation vector instead of a matrix. The CV Visual Odometry (Sequence) node does exactly this chain across a clip and flags frames where the estimate failed, which is a much better starting point than rebuilding it by hand.

Install

Manager → "ComfyUI CV", or:

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"

Needs Python ≥ 3.12 and a recent ComfyUI (V3 node API). Restart, then reload the browser page.

Traps

Configuration degeneracy. A purely rotating camera, or a scene that's one flat plane, gives you no usable pose - there's no baseline to measure. Use views with real depth and actual camera translation, and treat a tiny retval as "this pair failed", not "the pose is subtle".

Scale. Covered above, but it's the number one misunderstanding. t is a direction.

Which overload? This one for two calibrated cameras. If you already have an essential matrix, or you have a single shared K, the other three variants are simpler:

# the curated node does findEssentialMat + this decomposition + a found flag
# and is what the example workflows that do two-view SfM actually use

The pack's own note is that this wrapper is uncurated: no found flag, no failure tolerance, no shape validation. The curated CV Recover Pose (Essential Matrix) node returns found, an identity rotation and a zero translation instead of halting, which is what you want in a workflow that runs unattended.

Categoryimage/CV/low-level/cv2 R

Inputs (10)

NameTypeDefaultDescription
points1NPARRAYArray of N 2D points from the first image. The point coordinates should be floating-point (single or double precision). A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here.
points2NPARRAYArray of the second image points of the same size and format as points1 . A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here.
cameraMatrix1NPARRAYInput/output camera matrix for the first camera, the same as in A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here.
distCoeffs1NPARRAYInput/output vector of distortion coefficients, the same as in A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here.
cameraMatrix2NPARRAYInput/output camera matrix for the first camera, the same as in A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here.
distCoeffs2NPARRAYInput/output vector of distortion coefficients, the same as in A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here.
methodoptCOMBORANSACMethod for computing an essential matrix. - for the RANSAC algorithm. - for the LMedS algorithm.
proboptFLOAT0.9990-1e+38–1e+38Parameter used for the RANSAC or LMedS methods only. It specifies a desirable level of confidence (probability) that the estimated matrix is correct. Preset to the OpenCV default (0.999).
thresholdoptFLOAT1.0000-1e+38–1e+38Parameter used for RANSAC. It is the maximum distance from a point to an epipolar line in pixels, beyond which the point is considered an outlier and is not used for computing the final fundamental matrix. It can be set to something like 1-3, depending on the accuracy of the point localization, image resolution, and the image noise. Preset to the OpenCV default (1.0).
maskoptNPARRAY,IMAGE,MASKInput/output mask for inliers in points1 and points2. If it is not empty, then it marks inliers in points1 and points2 for the given essential matrix E. Only these inliers will be used to recover pose. In the output mask only inliers which pass the chirality check. This function decomposes an essential matrix using and then verifies possible pose hypotheses by doing chirality check. The chirality check means that the triangulated 3D points should have positive depth. Some details can be found in . This function can be used to process the output E and mask from . In this scenario, points1 and points2 are the same input for findEssentialMat.: Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.

Outputs (5)

NameTypeDescription
retvalINT—
ENPARRAY—
RNPARRAY—
tNPARRAY—
maskNPARRAY—