Nodes/ComfyUI CV/CV Solve PnP (Pose)
ComfyUI Node

CV Solve PnP (Pose)

Where is this chessboard in 3D space?

By bmad4ever·Created 4 months ago·Updated 15 days ago· 1
CV Solve PnP (Pose)
  • object_points
  • image_points
  • camera_matrix
  • dist_coeffs
  • rvec
  • tvec
  • found
  • inliers
  • inlier_count
  • reproj_error
◄methodSOLVEPNP_ITERATIVE►
◄estimationall points (least squares)►
◄reprojection_error3.0►
◄ransac_iterations1000►
◄confidence0.999►

Here's the problem PnP solves: you know the 3D positions of a set of points (a chessboard's corners, a cube's vertices, a marker's layout), you can see where those points land in the image, and you know your camera's intrinsics. From those three things, work out where the object is - rotation and translation.

That pair of vectors is what lets you draw a 3D overlay that sticks to a real object: axes on the chessboard, an AR frame on a marker, a wireframe cube sitting on a table. It's the classic computer-vision pose problem, and ComfyUI has no core node for it.

How it works

cv2.solvePnP under the hood, wrapped in the pack's failure-tolerant style. Fewer than four point correspondences, a count mismatch between the two point sets, or a degenerate configuration returns found = false with zero rvec/tvec instead of raising. Branch on the flag with a control-flow node and the overlay simply doesn't draw when detection failed - which is the behaviour you want on a batch where some frames have no board in them.

What you actually set

Four required inputs:

  • object_points - the 3D layout, Nx3 or Nx1x3. For a chessboard that's a planar grid from CV Grid Points; for a cube, hand-placed CV Points.
  • image_points - the matching 2D points, Nx2 or Nx1x2, in the same order as object_points. From cv2.findChessboardCorners, feature matches, or goodFeaturesToTrack. The order is not a detail you can get wrong and recover from: PnP assumes correspondence, so a shuffled list gives you a confidently wrong pose.
  • camera_matrix - the 3×3 intrinsics K, from CV Camera Matrix or a full Calibrate Camera run.
  • method - the solver. SOLVEPNP_ITERATIVE (Levenberg-Marquardt) is the general default; IPPE / IPPE_SQUARE are the ones for planar targets.

Then the optional ones, and the interesting choice is estimation:

  • all points (least squares) - fits every correspondence. Correct for a chessboard or a marker, where every point is trustworthy.
  • RANSAC - tolerates wrong correspondences, which is what feature matches or stereo-derived 3D points need. It also fills the inliers output. The RANSAC knobs are reprojection_error (3 px by default - how far a point may miss its reprojection and still count), ransac_iterations (1000) and confidence (0.999).

dist_coeffs is there for a real lens. Leave it unconnected for an ideal pinhole.

Outputs

rvec (3×1 Rodrigues rotation) and tvec (3×1 translation) are the pose - zeros when not found. found is your branch. inliers is an Nx1 uint8 mask of the correspondences RANSAC kept (all 255 in least-squares mode, all 0 when nothing was found); filter with CV Filter Points By Mask or CV Filter Points 3D By Mask. inlier_count is the count.

And reproj_error - the mean reprojection error in pixels over the inliers. This is the quality number, and the tooltip gives you the yardstick: under about 1–3 px is a good fit. Learn to read it before you trust a pose. A pose with a 40 px reprojection error will still draw axes; they'll just be wrong.

Feed rvec/tvec into cv2.projectPoints / cv2.drawFrameAxes to draw the overlay, into CV Pose To Matrix or CV ComposePose for the matrix form, or Rodrigues to convert the rotation representation.

Install

ComfyUI Manager → ComfyUI CV → install → restart. Or:

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"

Python ≥ 3.12, recent ComfyUI on the V3 node API. No models for the plain solver. The pack's example workflows that use PnP do pull in calibration sidecars - a Load Camera Params and the CV Install Example Inputs workflow that copies the sample calibration files into ComfyUI/input. Models are a separate matter entirely and are not bundled.

Common issues

  • found = false and you're sure the board was detected. Count mismatch between object_points and image_points is the usual cause - the corner detector found 54 corners and your grid declares 63. Check the two array shapes with CV CVArrayShape.
  • Axes drawn in the wrong place. Camera matrix from a different resolution than the image you're solving on. Intrinsics are resolution-specific; cv2.getOptimalNewCameraMatrix-style scaling is a real step, not an option.
  • Pose jitters frame to frame. You're in least-squares mode with sloppy correspondences. Switch estimation to RANSAC, watch reproj_error and inlier_count and set a threshold on them.
  • You want the pose ambiguity, not one answer. That's CV Solve PnP (All Poses) - the same problem, every solution.
  • Contrib nodes vanished after a pip install. Shared site-packages/cv2 overwritten by a non-contrib wheel. python tools/repair_opencv_contrib.py --check then --apply.
Categoryimage/CV/features

Inputs (9)

NameTypeDefaultDescription
object_pointsNPARRAY3D points in object space, Nx3 or Nx1x3 (e.g. 'CV Grid Points' for a chessboard, or 'CV Points' for a cube).
image_pointsNPARRAYThe matching 2D image points, Nx2 or Nx1x2, in the same order as object_points (e.g. findChessboardCorners).
camera_matrixNPARRAY3x3 intrinsic matrix K ('CV Camera Matrix' or 'Calibrate Camera').
methodCOMBOSOLVEPNP_ITERATIVEPnP solver. ITERATIVE (Levenberg-Marquardt) is the general default; IPPE/IPPE_SQUARE are for planar targets.
dist_coeffsoptNPARRAYLens distortion coefficients (k1, k2, p1, p2[, k3...]); leave unconnected for an ideal pinhole camera (no distortion).
estimationoptCOMBOall points (least squares)'all points' fits every correspondence - correct for a chessboard or marker, where they are all good. 'RANSAC' tolerates wrong correspondences and is what feature matches or stereo-derived 3D points need; it also fills the 'inliers' output.
reprojection_erroroptFLOAT3.00.1–100RANSAC only: how many pixels a point may miss its reprojection by and still count as an inlier.
ransac_iterationsoptINT10001–100000RANSAC only: maximum sampling iterations.
confidenceoptFLOAT0.9990–1RANSAC only: probability that the returned pose comes from an all-inlier sample.

Outputs (6)

NameTypeDescription
rvecNPARRAY3x1 Rodrigues rotation vector (zeros if not found). Pass to cv2.projectPoints / drawFrameAxes / Rodrigues / 'CV Pose To Matrix'.
tvecNPARRAY3x1 translation vector (zeros if not found).
foundBOOLEAN—
inliersNPARRAYNx1 uint8 mask of the correspondences RANSAC kept (all 255 for the 'all points' estimation, all 0 if not found). Filter the matches with 'CV Filter Points By Mask' / 'CV Filter Points 3D By Mask'.
inlier_countINT—
reproj_errorFLOATMean reprojection error in pixels over the inliers - the quality number for this pose (under ~1-3 px is a good fit). 0 when not found.