CV Solve PnP (Pose)
Where is this chessboard in 3D space?
- object_points
- image_points
- camera_matrix
- dist_coeffs
- rvec
- tvec
- found
- inliers
- inlier_count
- reproj_error
Here's the problem PnP solves: you know the 3D positions of a set of points (a chessboard's corners, a cube's vertices, a marker's layout), you can see where those points land in the image, and you know your camera's intrinsics. From those three things, work out where the object is - rotation and translation.
That pair of vectors is what lets you draw a 3D overlay that sticks to a real object: axes on the chessboard, an AR frame on a marker, a wireframe cube sitting on a table. It's the classic computer-vision pose problem, and ComfyUI has no core node for it.
How it works
cv2.solvePnP under the hood, wrapped in the pack's failure-tolerant style. Fewer than four point correspondences, a count mismatch between the two point sets, or a degenerate configuration returns found = false with zero rvec/tvec instead of raising. Branch on the flag with a control-flow node and the overlay simply doesn't draw when detection failed - which is the behaviour you want on a batch where some frames have no board in them.
What you actually set
Four required inputs:
- object_points - the 3D layout, Nx3 or Nx1x3. For a chessboard that's a planar grid from
CV Grid Points; for a cube, hand-placedCV Points. - image_points - the matching 2D points, Nx2 or Nx1x2, in the same order as object_points. From
cv2.findChessboardCorners, feature matches, orgoodFeaturesToTrack. The order is not a detail you can get wrong and recover from: PnP assumes correspondence, so a shuffled list gives you a confidently wrong pose. - camera_matrix - the 3×3 intrinsics K, from
CV Camera Matrixor a fullCalibrate Camerarun. - method - the solver.
SOLVEPNP_ITERATIVE(Levenberg-Marquardt) is the general default; IPPE / IPPE_SQUARE are the ones for planar targets.
Then the optional ones, and the interesting choice is estimation:
- all points (least squares) - fits every correspondence. Correct for a chessboard or a marker, where every point is trustworthy.
- RANSAC - tolerates wrong correspondences, which is what feature matches or stereo-derived 3D points need. It also fills the
inliersoutput. The RANSAC knobs arereprojection_error(3 px by default - how far a point may miss its reprojection and still count),ransac_iterations(1000) andconfidence(0.999).
dist_coeffs is there for a real lens. Leave it unconnected for an ideal pinhole.
Outputs
rvec (3×1 Rodrigues rotation) and tvec (3×1 translation) are the pose - zeros when not found. found is your branch. inliers is an Nx1 uint8 mask of the correspondences RANSAC kept (all 255 in least-squares mode, all 0 when nothing was found); filter with CV Filter Points By Mask or CV Filter Points 3D By Mask. inlier_count is the count.
And reproj_error - the mean reprojection error in pixels over the inliers. This is the quality number, and the tooltip gives you the yardstick: under about 1–3 px is a good fit. Learn to read it before you trust a pose. A pose with a 40 px reprojection error will still draw axes; they'll just be wrong.
Feed rvec/tvec into cv2.projectPoints / cv2.drawFrameAxes to draw the overlay, into CV Pose To Matrix or CV ComposePose for the matrix form, or Rodrigues to convert the rotation representation.
Install
ComfyUI Manager → ComfyUI CV → install → restart. Or:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"
Python ≥ 3.12, recent ComfyUI on the V3 node API. No models for the plain solver. The pack's example workflows that use PnP do pull in calibration sidecars - a Load Camera Params and the CV Install Example Inputs workflow that copies the sample calibration files into ComfyUI/input. Models are a separate matter entirely and are not bundled.
Common issues
found = falseand you're sure the board was detected. Count mismatch between object_points and image_points is the usual cause - the corner detector found 54 corners and your grid declares 63. Check the two array shapes withCV CVArrayShape.- Axes drawn in the wrong place. Camera matrix from a different resolution than the image you're solving on. Intrinsics are resolution-specific;
cv2.getOptimalNewCameraMatrix-style scaling is a real step, not an option. - Pose jitters frame to frame. You're in least-squares mode with sloppy correspondences. Switch
estimationto RANSAC, watchreproj_errorandinlier_countand set a threshold on them. - You want the pose ambiguity, not one answer. That's
CV Solve PnP (All Poses)- the same problem, every solution. - Contrib nodes vanished after a pip install. Shared
site-packages/cv2overwritten by a non-contrib wheel.python tools/repair_opencv_contrib.py --checkthen--apply.
Inputs (9)
| Name | Type | Default | Description |
|---|---|---|---|
| object_points | NPARRAY | 3D points in object space, Nx3 or Nx1x3 (e.g. 'CV Grid Points' for a chessboard, or 'CV Points' for a cube). | |
| image_points | NPARRAY | The matching 2D image points, Nx2 or Nx1x2, in the same order as object_points (e.g. findChessboardCorners). | |
| camera_matrix | NPARRAY | 3x3 intrinsic matrix K ('CV Camera Matrix' or 'Calibrate Camera'). | |
| method | COMBO | SOLVEPNP_ITERATIVE | PnP solver. ITERATIVE (Levenberg-Marquardt) is the general default; IPPE/IPPE_SQUARE are for planar targets. |
| dist_coeffsopt | NPARRAY | Lens distortion coefficients (k1, k2, p1, p2[, k3...]); leave unconnected for an ideal pinhole camera (no distortion). | |
| estimationopt | COMBO | all points (least squares) | 'all points' fits every correspondence - correct for a chessboard or marker, where they are all good. 'RANSAC' tolerates wrong correspondences and is what feature matches or stereo-derived 3D points need; it also fills the 'inliers' output. |
| reprojection_erroropt | FLOAT | 3.00.1–100 | RANSAC only: how many pixels a point may miss its reprojection by and still count as an inlier. |
| ransac_iterationsopt | INT | 10001–100000 | RANSAC only: maximum sampling iterations. |
| confidenceopt | FLOAT | 0.9990–1 | RANSAC only: probability that the returned pose comes from an all-inlier sample. |
Outputs (6)
| Name | Type | Description |
|---|---|---|
| rvec | NPARRAY | 3x1 Rodrigues rotation vector (zeros if not found). Pass to cv2.projectPoints / drawFrameAxes / Rodrigues / 'CV Pose To Matrix'. |
| tvec | NPARRAY | 3x1 translation vector (zeros if not found). |
| found | BOOLEAN | — |
| inliers | NPARRAY | Nx1 uint8 mask of the correspondences RANSAC kept (all 255 for the 'all points' estimation, all 0 if not found). Filter the matches with 'CV Filter Points By Mask' / 'CV Filter Points 3D By Mask'. |
| inlier_count | INT | — |
| reproj_error | FLOAT | Mean reprojection error in pixels over the inliers - the quality number for this pose (under ~1-3 px is a good fit). 0 when not found. |