cv2.rapid.rapid
One pose-refinement step of OpenCV's model tracker
- img
- pts3d
- tris
- K
- rvec
- tvec
- ratio
- rvec
- tvec
- rmsd
The 30-second version
You have a textured 3D mesh and a photo that contains it somewhere. cv2.rapid.rapid projects the mesh through the pose you give it, then nudges that pose once by matching the projected silhouette against image gradients, and hands back the fresh rotation and translation. Do it again per frame and you have a tracker.
It is a refinement step, and that word is load-bearing. RAPID (Harris & Stennett, 1990) has no detection stage: give it a pose that's off by a hundred pixels and it will happily converge on the wrong edges. It's for the case where you know roughly where the object is - from a marker, from the previous frame, from a hand-placed pose - and you want it to stay glued as the object moves.
It's a raw auto-generated wrapper over the contrib rapid submodule, so the widgets are cv2's own parameters in cv2's own order.
Inputs that matter
Eight required inputs, no optional ones, and every one of them is an array:
- img - the frame.
- num - how many control points to sample along the silhouette. The shipped example workflow passes 150. The pack's curated sequence node defaults to 128 and says below ~32 a couple of bad correspondences can swing the pose. The wrapper's own widget defaults to 0, which is not a usable setting - set it.
- len - the half-length of the search line in pixels: how far either side of the predicted contour the tracker looks. The example uses 30. Bigger buys capture range and costs you accuracy, because RAPID takes the strongest gradient anywhere on the line. The pack's sequence node measured search length 96 pinning the error at ~46 px on a textured subject over clutter, i.e. more rope doesn't help. Downscale instead.
- pts3d, tris - the mesh. Triangle winding matters: it drives backface culling and the silhouette. Build these with CV Mesh From 3D Model, and run the mesh through CV Mesh Split Long Edges if it's coarse - the pack warns a coarse mesh puts the control points tens of pixels off and the track diverges.
- K - 3×3 intrinsics. RAPID does not model lens distortion, so undistort wide-angle footage first.
- rvec, tvec - the starting pose, in the mesh's units.
Outputs: ratio (fraction of search lines that produced a usable correspondence - a relative health signal, bounded by 1), rvec and tvec (the refined pose), and rmsd (RAPID's own 2D reprojection difference, in pixels).
One ordering gotcha worth knowing: the pack's source notes say the binding's return order was measured, because the registry's generic kind names were misleading - the first value is the ratio (it stays ≤ 1) and the last is the rmsd (it grows with search length). So read ratio and rmsd as the pair they are: low rmsd with a wrong pose is entirely possible, since the tracker can settle on a self-consistent but shifted silhouette.
Wiring it up without losing your mind
The natural graph is a feedback loop - feed the new rvec/tvec back in for the next iteration, one frame at a time - and ComfyUI doesn't do loops that way. Two ways out:
- CV Rapid Track (Sequence) - the curated node that folds this over a whole IMAGE batch, warm-starting each frame from the last, and emits per-frame
rvecs/tvecsplusratios/rmsdsstacks. This is what you want for actual tracking. - The
CV Rapid Pose Refinesubgraph - the pack's own composition of extractControlPoints → extractLineBundle → findCorrespondencies → convertCorrespondencies → solvePnP, if you want to see the pieces.
Call the raw wrapper directly when you're doing something the sequence node doesn't: a manual iteration count in a static-shot experiment, a pose refine after your own detector, or a two-frame sanity test. To see what it's doing, wire the refined pose into cv2.projectPoints and draw with cv2.rapid.drawWireframe - a wireframe you can watch slide off the object is worth an hour of staring at rmsd.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"
Or grab "ComfyUI CV" from ComfyUI Manager and let it handle the dependency. Python ≥ 3.12 and a recent ComfyUI (V3 node API) are required. Restart after installing.
If it misbehaves
The node is missing entirely. RAPID is contrib-only, and a non-contrib wheel wipes the contrib submodules. Verify with python -c "import cv2; print(cv2.rapid.rapid)" and repair with tools/repair_opencv_contrib.py --check then --apply from the pack folder.
The pose jumps or the wireframe pins to a background edge. That's the search line doing what it says on the tin. Shorten len, shrink the frames before tracking (the sequence node's scale is the cheaper way to buy capture range), and check your mesh isn't too coarse for the silhouette.
Everything comes back as zero. You left num or len at the widget default of 0 - this node needs the example workflow's kind of numbers (150 and 30 are the shipped ones).
Inputs (8)
| Name | Type | Default | Description |
|---|---|---|---|
| img | NPARRAY,IMAGE,MASK | - - - Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| num | INT | 0-2147483648–2147483647 | - - - |
| len | INT | 0-2147483648–2147483647 | - - - |
| pts3d | NPARRAY,IMAGE,MASK | - - - Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| tris | NPARRAY,IMAGE,MASK | - - - Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| K | NPARRAY,IMAGE,MASK | - - - Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| rvec | NPARRAY,IMAGE,MASK | - - - Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| tvec | NPARRAY,IMAGE,MASK | - - - Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. |
Outputs (4)
| Name | Type | Description |
|---|---|---|
| ratio | FLOAT | — |
| rvec | NPARRAY | — |
| tvec | NPARRAY | — |
| rmsd | FLOAT | — |