Nodes/ComfyUI CV/cv2.solvePnPRefineVVS
ComfyUI Node

cv2.solvePnPRefineVVS

Pose Refinement With a Gain Knob

By bmad4ever·Created 3 months ago·Updated 14 days ago· 1
cv2.solvePnPRefineVVS
  • objectPoints
  • imagePoints
  • cameraMatrix
  • distCoeffs
  • rvec
  • tvec
  • rvec
  • tvec
◄criteria_typemax count or epsilon (whichever first)►
◄criteria_max_count30►
◄criteria_epsilon0.00►
◄VVSlambda1.0000►

This is the sibling of cv2.solvePnPRefineLM, and the two are close enough that you should pick one and get on with your day. Both take a pose you already have and iteratively shave the reprojection error; both need the same six inputs and return the same rvec/tvec pair. The difference is how they take steps: LM uses Levenberg-Marquardt (a damped Gauss-Newton scheme that adapts its own damping), VVS formulates the same minimisation as a virtual visual servoing control law and makes the gain explicit.

That gain is VVSlambda, and it's the reason to pick this node over LM: there's a knob. The default is 1.0, "equivalent to the α gain in the damped Gauss-Newton formulation" in OpenCV's own words - lower values take smaller, more conservative steps; higher values push harder and risk overshoot. If LM oscillates or stalls on your data, this is the variant with something to turn.

Inputs

Everything here is required - no optional sockets to skip:

  • objectPoints - Nx3 object-space points. OpenCV documents the VVS refiner as needing at least 3 points, which makes it the minimal-geometry option when you haven't got many correspondences.
  • imagePoints - the matching Nx2 image points, same order.
  • cameraMatrix - your 3×3 K, from CV Camera Matrix or a calibration node.
  • distCoeffs - the distortion vector. Required, unlike on cv2.solvePnP; a zero vector via CV Scalar stands in for an ideal pinhole.
  • rvec / tvec - the pose to refine. This is a local optimiser: the input values are used as the initial solution, so they need to be in the right basin already.
  • criteria_type / criteria_max_count / criteria_epsilon (optional) - the convergence settings, presented as a readable group (stop on iteration cap, on epsilon, or whichever comes first; defaults 30 iterations and 0.001).
  • VVSlambda (optional, default 1.0) - the gain.

Outputs

Two sockets: rvec and tvec, the refined pose. No success flag - the node assumes it improved something, and confirming that is on you.

Which one - LM or VVS?

Try LM first. It's the more common choice, it self-regulates, and it has fewer ways to be misconfigured. Move to VVS when LM's convergence is the problem: it's the variant you tune rather than debug. Lower VVSlambda if the refinement jumps around or diverges; leave it at 1.0 if you have no reason to move it. If neither converges from your starting pose, the starting pose is the problem, not the optimizer.

For a pose obtained from noisy correspondences, the full good-practice chain is: solvePnPRansac → (optionally index to just the inliers with CV Take By Index) → refine with LM or VVS → verify by reprojecting with cv2.projectPoints. The pack's CV Rapid Pose Refine subgraph follows that shape for mesh tracking, warm-starting each frame from the last refined pose.

Install

Manager → search ComfyUI CV, or:

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"

Restart ComfyUI. Python ≥ 3.12 and a recent ComfyUI on the V3 node API. No models, no downloads.

Where people get burned

Garbage in, polished garbage out. A local refiner cannot escape a bad initial pose, and it will report nothing when it fails. If your refined pose is wrong, the upstream solve is the suspect - check retval on solvePnPRansac, not VVSlambda.

Distortion required. Forgetting that distCoeffs has no "leave blank" option here is the most common first-run error on this node, especially if you came from cv2.solvePnP, where it's optional.

Outlier sensitivity. Like LM, VVS minimises over every point it's given. Pre-filter to the inliers on noisy data or one bad correspondence drags the fit.

Point correspondence order. objectPoints[i] must pair with imagePoints[i]. A shuffled pairing produces a tight, confident, wrong pose.

Contrib-wheel collisions. The four OpenCV distributions share one site-packages/cv2; installing a non-contrib wheel over the contrib build leaves contrib-backed nodes missing from the menu with no error. tools/repair_opencv_contrib.py --check diagnoses it and --apply repairs it. When a whole family of nodes "disappears" from this pack, that's almost always why.

Categoryimage/CV/low-level/cv2 S

Inputs (10)

NameTypeDefaultDescription
objectPointsNPARRAYArray of object points in the object coordinate space, Nx3 1-channel or 1xN/Nx1 3-channel, where N is the number of points. vector\ can also be passed here. A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here.
imagePointsNPARRAYArray of corresponding image points, Nx2 1-channel or 1xN/Nx1 2-channel, where N is the number of points. vector\ can also be passed here. A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here.
cameraMatrixNPARRAYInput camera intrinsic matrix $\cameramatrix{A}$ . A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here.
distCoeffsNPARRAYInput vector of distortion coefficients $\distcoeffs$. If the vector is NULL/empty, the zero distortion coefficients are assumed. A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here.
rvecNPARRAYInput/Output rotation vector (see ) that, together with tvec, brings points from the model coordinate system to the camera coordinate system. Input values are used as an initial solution. A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here.
tvecNPARRAYInput/Output translation vector. Input values are used as an initial solution. A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here.
criteria_typeoptCOMBOmax count or epsilon (whichever first)When to stop iterating: after max_count iterations, when the change drops below epsilon, or whichever comes first.
criteria_max_countoptINT301–2147483647Maximum iterations (ignored when 'epsilon only').
criteria_epsilonoptFLOAT0.000–1e+38Target accuracy / smallest change worth continuing for (ignored when 'max count only').
VVSlambdaoptFLOAT1.0000-1e+38–1e+38Gain for the virtual visual servoing control law, equivalent to the $\alpha$ gain in the Damped Gauss-Newton formulation. The function refines the object pose given at least 3 object points, their corresponding image projections, an initial solution for the rotation and translation vector, as well as the camera intrinsic matrix and the distortion coefficients. The function minimizes the projection error with respect to the rotation and the translation vectors, using a virtual visual servoing (VVS) scheme. Preset to the OpenCV default (1.0).

Outputs (2)

NameTypeDescription
rvecNPARRAY—
tvecNPARRAY—