Nodes/ComfyUI CV/CV Solve PnP (All Poses)
ComfyUI Node

CV Solve PnP (All Poses)

The pose problem has more than one answer

By bmad4ever·Created 4 months ago·Updated 15 days ago· 1
CV Solve PnP (All Poses)
  • object_points
  • image_points
  • camera_matrix
  • dist_coeffs
  • rvecs
  • tvecs
  • reprojection_error
  • count
  • best_index
  • found
◄solvergeneric - every solution the method finds►
◄methodSOLVEPNP_ITERATIVE►

Beginners assume pose estimation has one right answer that a solver finds. Anyone who's done it knows the geometry is often ambiguous: three point correspondences admit up to four poses, and a planar target viewed near head-on has a well-known two-fold flip - the board leaning toward you and away from you project almost identically. cv2.solvePnP picks one and hands it over with no hint that a sibling exists.

CV Solve PnP (All Poses) returns the whole stack instead, each candidate with its own reprojection error, so you can see whether the answer was obvious or a coin flip.

How it works

cv2.solveP3P / cv2.solvePnPGeneric under the hood, through a solver dropdown that decides the shape of the call:

  • generic - every solution the method finds - runs whatever method you select in the second dropdown over all the points (4 minimum). SOLVEPNP_ITERATIVE returns a single refined pose; IPPE / IPPE_SQUARE are for planar targets and are the ones that expose the two-fold ambiguity; SQPNP and EPNP are the non-iterative alternatives.
  • P3P / AP3P - take exactly three or four points and return up to four poses. With four points the extra one only ranks the candidates. AP3P is the newer formulation and better behaved numerically, which is the version to use if you have the choice.

Same failure tolerance as the rest of the pack: too few points, a count mismatch or a degenerate configuration gives found = false with empty stacks rather than an exception.

What you set

  • object_points - Nx3 or Nx1x3, same as the single-pose node. CV Grid Points for a planar target, CV Points otherwise.
  • image_points - the matching Nx2 or Nx1x2, in the same order.
  • camera_matrix - the 3×3 K.
  • solver and method - the two dropdowns above. Note that method is ignored by P3P/AP3P; it only steers the generic solver.
  • dist_coeffs - optional, for a real lens.

The outputs, and how to read them

  • rvecs - (N, 3) float64, one Rodrigues vector per candidate. tvecs - (N, 3), same order, in the units of your object points.
  • reprojection_error - (N, 1) float64, mean reprojection error in pixels for each candidate, computed identically for every solver (a cv2.projectPoints pass over all the points). Under about 1 px is a good fit.
  • count - how many candidates came back, 0 when not found.
  • best_index - the index of the lowest-error candidate, or -1 when not found.
  • found - BOOLEAN.

The gap between best and second-best is the number that's actually interesting, and the tooltip says so: a large gap means the ambiguity was resolved by the data; a small one means both poses fit your correspondences about equally well and you should not be reading anything into which one won. That's the whole reason to use this node over its single-answer sibling.

Take a specific candidate with CV Index Batch using best_index (or any index you want to inspect), then feed it to CV Pose To Matrix / cv2.projectPoints / drawFrameAxes and draw it. Drawing the top two candidates side by side is the fastest way to see what's ambiguous about your setup - and the fix is usually geometric, not algorithmic: get more points off the plane, or take the shot from a less head-on angle.

Install

ComfyUI Manager → ComfyUI CV → install → restart. Manual:

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"

Python ≥ 3.12, recent ComfyUI on the V3 node API, no models required for the solver. One install note that's specific to this pack: some of its example workflows and subgraphs assume ComfyUI-Inspire-Pack, ComfyUI-Custom-Scripts or Basic Data Handling for loops and list handling. This node doesn't, but if a graph you copied off the repo opens with red nodes, that's usually why.

Common issues

  • count = 1 when you expected four. Normal. SOLVEPNP_ITERATIVE refines to a single answer; only the planar methods and P3P families fan out. If you want the ambiguity you have to pick a solver that exposes it.
  • found = false with P3P and five points. P3P takes exactly 3 or 4 points and the count mismatch is reported as not-found rather than raising. Trim the arrays, or use the generic solver.
  • Two candidates with nearly equal error and you picked one. You didn't resolve anything, you coin-flipped. Add correspondences off the target plane, or gate on the error gap before acting on a pose.
  • Poses on a planar grid look like a mirror pair. That's the classic planar ambiguity, showing up exactly as documented. Use AP3P or IPPE_SQUARE deliberately, and inspect both candidates.
  • All the CV nodes disappeared after another pack's install. A non-contrib OpenCV wheel over the contrib one - shared site-packages/cv2, contrib submodules emptied. python tools/repair_opencv_contrib.py --check, then --apply.
Categoryimage/CV/features

Inputs (6)

NameTypeDefaultDescription
object_pointsNPARRAY3D points in object space, Nx3 or Nx1x3 ('CV Grid Points' for a planar target, 'CV Points' otherwise).
image_pointsNPARRAYThe matching 2D image points, Nx2 or Nx1x2, in the same order as object_points.
camera_matrixNPARRAY3x3 intrinsic matrix K ('CV Camera Matrix' or 'Calibrate Camera').
solverCOMBOgeneric - every solution the method finds'generic' runs the solver named below over ALL the points (4 minimum). P3P/AP3P take EXACTLY 3 or 4 points and return up to four poses - with 4 the extra point only ranks them. AP3P is the newer, numerically better-behaved formulation.
methodCOMBOSOLVEPNP_ITERATIVEWhich algorithm the 'generic' solver runs (ignored by P3P/AP3P). ITERATIVE returns a single refined pose; IPPE / IPPE_SQUARE are for PLANAR targets and are the ones that expose the two-fold ambiguity; SQPNP and EPNP are non-iterative alternatives.
dist_coeffsoptNPARRAYLens distortion coefficients (k1, k2, p1, p2[, k3...]); leave unconnected for an ideal pinhole camera.

Outputs (6)

NameTypeDescription
rvecsNPARRAY(N, 3) float64 stack of Rodrigues rotation vectors, one row per candidate pose. Take one with 'CV Index Batch' and feed 'CV Pose To Matrix' / cv2.projectPoints / drawFrameAxes.
tvecsNPARRAY(N, 3) float64 stack of translation vectors, in the same order (and units) as object_points.
reprojection_errorNPARRAY(N, 1) float64 mean reprojection error in PIXELS for each candidate, computed the same way for every solver (cv2.projectPoints over all the points). Under ~1 px is a good fit; a large gap between the best and the second best means the ambiguity is resolved.
countINTHow many candidate poses were returned (0 when not found).
best_indexINTIndex of the candidate with the lowest reprojection error, or -1 when not found.
foundBOOLEAN—