Nodes/ComfyUI CV/cv2.rapid.extractControlPoints
ComfyUI Node

cv2.rapid.extractControlPoints

Where RAPID's silhouette samples come from

By bmad4ever·Created 4 months ago·Updated 15 days ago· 1
cv2.rapid.extractControlPoints
  • pts3d
  • rvec
  • tvec
  • K
  • imsize
  • tris
  • ctl2d
  • ctl3d
◄num0►
◄len0►

RAPID tracks a 3D model in an image by putting points along the model's projected silhouette and searching outward from each one for a real image edge. This node is the "put the points" step, and it hands back both halves of the pairing you'll need later: the pixel positions of the samples and the 3D points those pixels correspond to.

What it produces

Two NPARRAY outputs:

  • ctl2d - the sampled control points in the image: pixel positions along the projected model's silhouette.
  • ctl3d - the matching object-space points.

The pack's own return documentation notes the pair order follows the measured Python binding rather than the C++ parameter order - a small reminder that this is an auto-generated wrapper over a contrib module, not a papered-over API. The practical upshot: don't assume which output is which from the ordering you read in the C++ docs. Check shapes once with an inspection node, then wire accordingly.

The 3D half is the one you keep. ctl3d gets paired with the image points that the search eventually lands on (from rapid.convertCorrespondencies' pts2d) and handed to a PnP solve. The pack's return docs say this explicitly: pts2d comes back aligned with extractControlPoints' ctl3d, ready for cv2.solvePnP. That alignment is the contract of the whole pipeline, so if the two arrays are resampled, filtered or reordered between the stages, nothing downstream makes sense.

Inputs, and what each one is for

  • num - how many control points to sample. More points means a more constrained pose but a slower round and more chances to lock onto clutter. Treat it as your main quality/cost dial.
  • len - the half-length of the search lines in pixels. It's an input here because the search lines are constructed together with the samples; too short and the line never reaches the edge, too long and it finds background clutter.
  • pts3d - the model's vertices as an N×3 cloud. From CV Mesh From 3D Model (a loaded .glb/.obj), from a PLY you loaded, or from CV Transform Points 3D if you've already moved it into the pose's frame. For RAPID, this is the mesh.
  • rvec, tvec - the current pose. This is a refinement step: the incoming pose has to be close, or the projected silhouette won't be near the real edges and the samples will be in the wrong place entirely.
  • K - the 3×3 intrinsics. CV Camera Matrix, a calibration result, or CV Load Camera Params (JSON).
  • imsize - the image size as a CV_TUPLE, (width, height), typed in place or wired from CV Tuple.
  • tris - the triangle index array (M×3 integers) describing the mesh's faces, from CV Mesh From 3D Model.

Note the shape of this: eight inputs, all data, no pictures. Despite the pack's polymorphic sockets accepting an IMAGE link on these, wiring one in is a mistake - cv2 will not interpret a photograph as a mesh.

Where it sits

The full RAPID round in this pack's raw nodes:

  1. rapid.extractControlPoints - this node: samples + their 3D partners.
  2. rapid.extractLineBundle - collect image intensities along each sample's normal, plus the pixel locations of those samples.
  3. rapid.findCorrespondencies and rapid.convertCorrespondencies - find the best edge along each line and turn it into pts2d plus a validity mask.
  4. rapid.rapid - the whole iteration in one call, or a PnP solve you assemble yourself.
  5. Debug with rapid.drawSearchLines / rapid.drawCorrespondencies / rapid.drawWireframe when it drifts.

For the packaged version, the CV Rapid Pose Refine subgraph exposes exactly this as a single node, and the curated CV Rapid Track (Sequence) runs it across a batch with warm-started poses.

Install

ComfyUI CV by bmad4ever. ComfyUI Manager, search comfyui_cv, or:

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"

Restart. Python ≥ 3.12, recent V3-API ComfyUI. For any practical use, remember the author's own disclaimer: the example workflows are showcases, not production pipelines.

Common issues

Node absent. cv2.rapid is contrib-only; a non-contrib OpenCV wheel removes the submodule and the pack skips it. tools/repair_opencv_contrib.py --check / --apply.

Control points land nowhere near the object. The incoming pose is too far off, or K doesn't match the image resolution. RAPID refines; it doesn't search.

The pose inverts. Normals again. If your cloud came from a mesh, use CV Mesh Vertex Normals - a plane fit can't orient a closed model and will flip the pose by roughly 180°.

"Mesh is too dense, RAPID is slow." CV Mesh Split Long Edges gives you denser control-point coverage at a fraction of the triangles, which the pack recommends over uniform subdivision for exactly this reason. And if you're sitting on a genuine mesh, remember from 3d-generation.md that generated meshes are triangle soup - the topology you feed RAPID is usually worse than the topology you'd hand a human.

Categoryimage/CV/low-level/rapid

Inputs (8)

NameTypeDefaultDescription
numINT0-2147483648–2147483647 - - -
lenINT0-2147483648–2147483647 - - -
pts3dNPARRAY,IMAGE,MASK - - - Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
rvecNPARRAY,IMAGE,MASK - - - Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
tvecNPARRAY,IMAGE,MASK - - - Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
KNPARRAY,IMAGE,MASK - - - Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
imsizeCV_TUPLE0,0One value with 2 components (w, h) - it travels as a whole, so it cannot arrive half-connected. Wire it from 'CV Tuple' or type the components in place.
trisNPARRAY,IMAGE,MASK - - - Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.

Outputs (2)

NameTypeDescription
ctl2dNPARRAY—
ctl3dNPARRAY—