cv2.solvePnP
Where Is That Object, Actually? (Pose From Points)
- objectPoints
- imagePoints
- cameraMatrix
- distCoeffs
- retval
- rvec
- tvec
solvePnP is the operation that turns a camera into a measuring instrument. You know where a set of 3D points are in the world (a board, a marker, a model), you've found where they appear in the image, and you know your camera's intrinsics - so the only unknown is where the camera is relative to the object. Perspective-n-Point solves for that: a rotation vector and a translation vector, the classic rvec/tvec pair that drives every AR overlay.
In ComfyUI CV (bmad4ever/comfyui_cv) this is the raw wrapper. It's the same math as the pack's curated CV Solve PnP (Pose) node, which adds failure tolerance (found) and a cleaner interface - use the curated one if you just want a pose. Reach for the raw wrapper when you want the exact OpenCV behaviour, or you're rebuilding the pipeline by hand.
Inputs
objectPoints- the 3D points in object space,Nx3(orNx1/1xNwith 3 channels). CV Grid Points builds a planar grid from a chessboard pattern; CV Points lets you type a cube's corners.imagePoints- the matching 2D points,Nx2, in the same order. From cv2.findChessboardCorners, the ArUco detection nodes, a matcher's output, orcv2.goodFeaturesToTrack.cameraMatrix- the 3×3 intrinsic matrixK. Build it in one node with CV Camera Matrix (fx, fy, cx, cy), or get a real one out of the calibration nodes.distCoeffs(optional) - the lens distortion vector. Leave it unconnected for an ideal pinhole; that path is supported and means zero distortion.useExtrinsicGuess(optional, default false) - for the iterative solver, use provided rvec/tvec as a starting approximation. Leave it off on this node:rvec/tvecare outputs here, not inputs, so there is no guess to seed with. It's in the signature because OpenCV has it.flags(optional, defaultSOLVEPNP_ITERATIVE) - the solver.ITERATIVE(Levenberg-Marquardt) is the general default;EPNPandSQPNPare the non-planar/global options;IPPEandIPPE_SQUAREare for planar targets and are the right pick for a chessboard or a flat marker;P3PandAP3Pare the minimal-point solvers, and OpenCV wants exactly four points for those rather than the three their name suggests.
Outputs
Three sockets, named after cv2's own return values:
retval- a boolean.falsemeans no pose. Do not ignore it.rvec- the rotation vector (axis-angle).cv2.Rodriguesturns it into a matrix; CV Pose To Matrix gives you the 4×4.tvec- the translation vector.
From there the useful downstream moves are cv2.projectPoints to verify the pose by re-projecting the object points, cv2.drawFrameAxes to draw the coordinate frame on the frame (the pack's 57_camera_pose_3d_preview.json does exactly this and then loads a 3D model into the pose with the viewer nodes), or CV Compose Pose (Trajectory) to chain several poses into a trajectory.
Install
Manager → search ComfyUI CV → install, or:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"
Restart afterwards. Python ≥ 3.12 and a recent V3-API ComfyUI. No model downloads - this is pure geometry, though the example workflows do need the pack's sample inputs copied into ComfyUI/input (run workflows/01_install_example_inputs.json, then reload the page).
Where people get burned
Point order is the whole game. objectPoints[i] must correspond to imagePoints[i]. A shuffled order still returns a pose - a wrong one that reprojects plausibly. Chessboard detection returns corners in a defined row-major order, which is why the calibration flows work; hand-built point sets are where this bites.
Point units. The translation vector comes back in whatever units the object points were in. Millimetres in, millimetres out - and the pose will look completely reasonable either way, which is why scale errors survive so long.
Planarity and point count. Planar targets want IPPE/IPPE_SQUARE; P3P/AP3P want exactly four points; the iterative solver wants a well-conditioned set. Feeding the wrong solver the wrong geometry is the second-most-common cause of a plausible-but-wrong answer.
Trusting rvec when retval is false, or when you used RANSAC. For noisy correspondences, cv2.solvePnPRansac is the tolerant version, and it also hands you the inlier indices. For a pose that's close but not converged, a refinement pass is the third step (see cv2.solvePnPRefineLM, which the pack's CV Rapid Pose Refine subgraph uses).
Contrib-wheel collisions. All four OpenCV distributions install into the same site-packages/cv2; putting a non-contrib wheel over the contrib build leaves contrib-backed nodes missing from the menu with no error. tools/repair_opencv_contrib.py --check diagnoses it.
Inputs (6)
| Name | Type | Default | Description |
|---|---|---|---|
| objectPoints | NPARRAY | Array of object points in the object coordinate space, Nx3 1-channel or 1xN/Nx1 3-channel, where N is the number of points. vector\ can be also passed here. A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here. | |
| imagePoints | NPARRAY | Array of corresponding image points, Nx2 1-channel or 1xN/Nx1 2-channel, where N is the number of points. vector\ can be also passed here. A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here. | |
| cameraMatrix | NPARRAY | Input camera intrinsic matrix $\cameramatrix{A}$ . A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here. | |
| distCoeffsopt | NPARRAY | Input vector of distortion coefficients $\distcoeffs$. If the vector is NULL/empty, the zero distortion coefficients are assumed. Optional - leave unconnected for the OpenCV default (None). A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here. | |
| useExtrinsicGuessopt | BOOLEAN | false | Parameter used for #SOLVEPNP_ITERATIVE. If true (1), the function uses the provided rvec and tvec values as initial approximations of the rotation and translation vectors, respectively, and further optimizes them. Preset to the OpenCV default (False). |
| flagsopt | COMBO | SOLVEPNP_ITERATIVE | Method for solving a PnP problem: see More information about Perspective-n-Points is described in |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| retval | BOOLEAN | — |
| rvec | NPARRAY | — |
| tvec | NPARRAY | — |