Nodes/ComfyUI CV/CV Recover Pose (Essential Matrix)
ComfyUI Node

CV Recover Pose (Essential Matrix)

Monocular geometry gives you a direction, never a distance

By bmad4ever·Created 4 months ago·Updated 15 days ago· 1
CV Recover Pose (Essential Matrix)
  • points_a
  • points_b
  • camera_matrix
  • rotation
  • translation
  • inlier_mask
  • inlier_count
  • found
◄methodRANSAC►
◄threshold1.0►
◄confidence0.999►
◄min_inliers15►

This is the core step of two-view structure-from-motion, and the whole thing rests on one fact you should internalise before wiring it: it hands you the camera's rotation and the direction of its translation. The magnitude is unknowable from two images alone. A coin and a mountain look identical. If you skip that, you'll spend an afternoon wondering why your trajectory has a scale of 1.

What it does

Given matched points between two views of a static scene and the camera's intrinsics, it estimates the essential matrix (cv2.findEssentialMat with RANSAC) and decomposes it (cv2.recoverPose), running the cheirality check: of the four possible decompositions, it keeps the one where the matched points land in front of both cameras. That check is what makes the answer physically meaningful rather than algebraically valid.

It's the packaged version of a step that's genuinely tedious to do by hand, and it can't be done with the pack's raw wrappers at all - those are generated from top-level cv2 functions, and this needs findEssentialMat and recoverPose composed with a shape and convention bridge in between.

Inputs

points_a and points_b are Nx1x2 matched points from CV Match Features, in view 1 and view 2 respectively. camera_matrix is the 3x3 K shared by both views - CV Camera Matrix for a synthetic or already-calibrated rig, CV Calibrate Camera (Chessboard) when you've gone and calibrated. If you don't know your intrinsics, this node is not the place to hand-wave.

method defaults to RANSAC, which is what you want for feature matches. threshold (1 px) is the epipolar inlier distance; confidence (0.999) is the RANSAC confidence. min_inliers (15) is this pack's own addition and the one worth understanding: it's the minimum number of cheirality-consistent inliers before found goes true. Raise it to reject accidental fits - a good match set will clear 15 easily, and a bad one will hover around it forever.

Outputs

rotation is a 3x3 float64 R mapping view-1 directions into view 2, translation is a 3x1 unit vector (camera-2's origin in view-1's frame, up to scale). inlier_mask is the per-point Nx1 uint8 that survived - feed it into CV Filter Points By Mask or CV Draw Matches to see what the solver used, which is the fastest way to diagnose a bad result. inlier_count and found finish it.

On failure - fewer than 5 points, a degenerate configuration, or too few inliers - you get found=false, an identity rotation and a zero translation rather than an exception. Branch on found with a control-flow node. The consistent contract across this pack is that a failed geometric estimate never kills the run, which is great for batch work and terrible if you're not checking the flag.

The degeneracies that will get you

Two cases where the math cannot work, and both are common in practice:

  • Pure rotation. Pan a camera on a tripod and there is no translation to recover. The scene must have depth and the camera must move.
  • A single plane. Feature matches all lying on one wall or one facade are degenerate for the essential matrix. You need 3-D structure in view.

Also worth remembering for downstream: the output is a relative pose for one pair. Turning a sequence of these into a trajectory means composing them, which is a different job with a different node in this pack (CV Compose Pose), and the pack's own notes are candid that a failed step should hold the previous pose rather than fabricate motion.

Install

Part of ComfyUI CV (bmad4ever/comfyui_cv):

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"
# restart ComfyUI

Manager: search the pack title. Python ≥ 3.12, V3-API ComfyUI. The contrib wheel matters if you also want the feature detectors and matchers in the same pack - a non-contrib opencv-python installed over a contrib one silently empties the contrib submodules.

What to wire it to

The linear next step is CV Triangulate Points (Two-View) with the same point pairs and K, giving you a sparse 3-D sketch. The pack's workflows/exercise_two_view_sfm.json runs the full chain - ORB detect, match with the ratio test, this node, triangulate - and its notes are the reminder that matters: views need depth and a translated camera. If the goal is a dense cloud rather than a sparse one, that's the stereo path (CV Quasi-Dense Stereo or SGBM plus cv2.reprojectImageTo3D), which needs two cameras and a baseline - different problem, different pipeline.

Categoryimage/CV/features

Inputs (7)

NameTypeDefaultDescription
points_aNPARRAYPoints in view 1, Nx1x2 (from 'CV Match Features').
points_bNPARRAYCorresponding points in view 2, Nx1x2.
camera_matrixNPARRAY3x3 intrinsic matrix K shared by both views ('CV Camera Matrix' or 'Calibrate Camera').
methodCOMBORANSACRobust estimator for the essential matrix. RANSAC is the standard choice for feature matches.
thresholdFLOAT1.00.1–100Max distance in pixels from a point to its epipolar line to count as an inlier (RANSAC).
confidenceFLOAT0.9990.5–1Desired probability that the estimate is correct (RANSAC/LMedS).
min_inliersINT155–10000Minimum cheirality-consistent inliers for 'found' to be true. Raise it to reject accidental fits.

Outputs (5)

NameTypeDescription
rotationNPARRAY3x3 float64 rotation R mapping view-1 directions into view 2 (identity if not found).
translationNPARRAY3x1 float64 UNIT translation (camera-2 origin in view-1 frame, up to scale; zero if not found).
inlier_maskNPARRAYNx1 uint8: 1 = point used in the recovered pose. Feed into 'CV Filter Points By Mask' / 'Draw Matches'.
inlier_countINT—
foundBOOLEAN—