cv2.recoverPose (2/4)
Decompose an essential matrix you already have
- E
- points1
- points2
- cameraMatrix
- mask
- retval
- R
- t
- mask
Why this variant exists
The four cv2.recoverPose overloads differ in what they expect you to already know. Variant 1/4 is the one for when the answer to "what do I already know?" is the essential matrix: recoverPose(E, points1, points2, cameraMatrix).
That matters because estimating E and decomposing E are separate jobs, and the first one is where you make choices. cv2.findEssentialMat has its own RANSAC settings, its own probability, its own threshold; maybe you filtered correspondences beforehand; maybe you want to reuse one E against several point sets. The full-pipeline overload (1/4) does the estimation for you with its own defaults. This one doesn't - it takes your E and returns the pose.
Everything else is the same physics: an essential matrix decomposes into four candidate (R, t) pairs, and cv2 keeps the one where the triangulated points land in front of both cameras - the cheirality check. So this node is doing a genuinely useful thing with your matrix, not just unpacking it.
Inputs and outputs
Required (all NPARRAY, no image sockets - these are arrays of numbers, and the brief is explicit that only an NPARRAY link is accepted):
- E - 3×3 essential matrix, from the cv2.findEssentialMat wrapper, or from the curated CV Recover Pose (Essential Matrix) node, which runs the estimator for you.
- points1, points2 - the same matched point sets you gave the estimator. Nx2, float. The chirality check needs them; the node cannot infer depth from nothing.
- cameraMatrix - 3×3 intrinsics shared by both views.
Optional: mask - an inlier mask for the points. If you have one from the estimation step, pass it: the cheirality check is then only evaluated on those points.
Outputs: retval (number of inliers used), R (3×3 rotation from view 1 to view 2), t (3×1 translation direction - unit scale, see below), mask (N×1 uint8, inliers that also passed chirality).
Note there's no E output here: you supplied it. That's the whole point of the variant.
The scale thing, again
t is a unit vector. Monocular geometry constrains direction only - nothing in the image tells you whether the camera moved 10 cm or 10 m. Chain these poses into a trajectory and the trajectory is in arbitrary units, scaled to the first baseline. Anything metric has to come from elsewhere: a stereo baseline, a known object dimension, an IMU, or a similarity alignment afterwards (which is what CV Trajectory Error (ATE/RPE) does before scoring one trajectory against another). If you ignore this, your "drift in metres" number is fiction.
When to reach for it, and when not to
Reach for it when you want the estimation stage under your control - a custom mask, a fixed E for a batch of pairs, or an E that came from a curated node you already trust.
Don't reach for it if you just want the pose from two views and matched features, because the pack already ships that as a single node: CV Recover Pose (Essential Matrix) takes the points, K, a method, a threshold, a confidence and a min_inliers, runs findEssentialMat + this decomposition internally, and - importantly - returns a found boolean instead of throwing when the estimate collapses (fewer than 5 points, a degenerate configuration, too few inliers). Raw wrappers don't have that contract: they surface cv2's exception and stop the workflow. For a batch of a hundred frames where three pairs are garbage, the curated node is the difference between a run and a traceback.
The pragmatic split: use the curated node for pipelines and this wrapper when you're experimenting with the estimator itself and want to see how a hand-tuned E behaves.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"
ComfyUI Manager can do both if you search "ComfyUI CV". Needs Python ≥ 3.12 and a recent ComfyUI on the V3 node API - this pack's nodes are built with comfy_api.latest, so on an older ComfyUI they won't register at all. Restart after install.
Traps
Passing a fundamental matrix. F and E are easy to confuse and the failure is silent-ish: F = K⁻ᵀ E K⁻¹, and feeding an F here gives you a garbage-but-plausible pose. If your matrices came from cv2.findFundamentalMat, either convert properly or use the uncalibrated route (variant 3/4, which takes focal length and principal point instead of K).
Mismatched point order. points1[i] must correspond to points2[i], same count, same order. A mask from a different run is worse than no mask.
Shape conventions. The pack passes these straight through to cv2, so feed (N,1,2) or (N,2) - the cheap way to tell which shape you have is the pack's own CV Array Shape / CV CV Inspect node.
Inputs (5)
| Name | Type | Default | Description |
|---|---|---|---|
| E | NPARRAY | The output essential matrix. A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here. | |
| points1 | NPARRAY | Array of N 2D points from the first image. The point coordinates should be floating-point (single or double precision). A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here. | |
| points2 | NPARRAY | Array of the second image points of the same size and format as points1 . A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here. | |
| cameraMatrix | NPARRAY | Camera intrinsic matrix $\cameramatrix{A}$ . Note that this function assumes that points1 and points2 are feature points from cameras with the same camera intrinsic matrix. A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here. | |
| maskopt | NPARRAY,IMAGE,MASK | Input/output mask for inliers in points1 and points2. If it is not empty, then it marks inliers in points1 and points2 for the given essential matrix E. Only these inliers will be used to recover pose. In the output mask only inliers which pass the chirality check. This function decomposes an essential matrix using and then verifies possible pose hypotheses by doing chirality check. The chirality check means that the triangulated 3D points should have positive depth. Some details can be found in . This function can be used to process the output E and mask from . In this scenario, points1 and points2 are the same input for findEssentialMat.: Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. |
Outputs (4)
| Name | Type | Description |
|---|---|---|
| retval | INT | — |
| R | NPARRAY | — |
| t | NPARRAY | — |
| mask | NPARRAY | — |