CV Unproject Points
Turning a click on a pixel into a 3D point
- points_2d
- depths
- K
- dist_coeffs
- points_3d
- valid
- count
A pixel is a ray, not a point. (437, 212) in an image is every location along one line out of the camera, and the only thing that collapses it to a single 3D position is knowing how far away that pixel is. CV Unproject Points takes pixel positions plus depth and hands you 3D points in camera space.
It's the inverse of cv2 projectPoints, for the case where you already know the depth. And unlike CV Depth to 3D Points, which does this for a whole depth image, this does it for the handful of points you actually care about - the pixel someone clicked, a bounding-box centre, a contour vertex, the corner of a detected object.
How it works
For each point it computes the camera-space ray from the pixel position and the 3x3 camera matrix, scales that by the depth, and you're done. Optionally, if you wire dist_coeffs, the points get undistorted first via cv2.undistortPoints - which is what you want for pixels measured on a real photograph, and what you leave unwired for a synthetic pinhole render.
The convention that matters, and the node says it plainly: depth is camera-space Z, the distance along the optical axis - not the distance from the camera centre. Those two are the same in the middle of the frame and diverge toward the edges. Feed it a distance-to-camera (a range image, some lidar-derived maps, a ray-length depth) and everything off-centre lands too far out, progressively worse as you approach the frame edges. It's the single most common way to get plausible-but-wrong points out of this node. It's the same convention CV Rasterize Mesh uses, so renders and depth maps in the pack agree with each other.
Inputs and outputs
- points_2d -
Nx2(orNx1x2) pixel positions.CV Annotate Pointsproduces these from clicks on an image, which is the workflow this node is built for. - depths - N values, one per point. Sample them from a depth map with
CV Sample Array At Points, or pass a single value, which broadcasts to every point - that's how you unproject onto a plane at a known distance. - K - the
3x3camera matrix the points were seen through. Must be the same camera that produced the depth, obviously, and a guess here bends everything. - dist_coeffs - optional lens distortion.
Outputs are points_3d (Nx3 float32, camera space: X right, Y down, Z into the scene), valid (N uint8, 0 where the depth was missing), and count. Points with a non-finite or ≤ 0 depth come back as NaN rather than raising, so a click that landed on the background is still a valid row you can filter - and valid is how you filter it.
Camera space isn't usually where you want to end up. To get into the model's or world's frame, invert the pose that rendered the image: OpenCV Pose To Matrix → cv2_invert → CV Transform Points 3D. That's exactly the chain the pack's CV Pick Point On Mesh blueprint wires up, and it's the least obvious part of the whole pipeline if you've never done it: the pose maps model→camera, so you need its inverse to go camera→model.
Installing it
Manager → ComfyUI CV, or:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
Restart. Python ≥ 3.12, current ComfyUI on the V3 node API, opencv-contrib-python-headless~=5.0.0.93. No model downloads. GPL-3.0, a fork of opencv-comfyui, LLM-assisted code - the author's own recommendation is to read it before relying on it, which for a two-line unprojection is cheap to do.
Where people get burned
Depth from a model, not from geometry. If your "depth" came out of a monocular depth estimator, it's relative and up to an unknown scale - so your 3D points are correct in shape and arbitrary in size. Good for placing a marker on an object in a render, useless for measuring anything in metres until you calibrate the scale.
Mixing up K and the render. A camera matrix with a different focal length than the one that made the image gives you points on a plausible-looking arc that doesn't sit on the surface. Overlay them on the render to check before you build anything on top.
NaN rows slipping downstream. NaNs propagate silently through arithmetic and will poison a mean, a fit, or an ICP registration. Gate on valid or filter the array first.
Assuming a click is a single point. It's a pixel - so it's the centre of a pixel's ray. For placing an anchor that's irrelevant; for precise measurements it's the same sub-pixel argument as everywhere else.
Inputs (4)
| Name | Type | Default | Description |
|---|---|---|---|
| points_2d | NPARRAY | Nx2 (or Nx1x2) pixel positions. 'CV Annotate Points' produces these from clicks on the image. | |
| depths | NPARRAY | N camera-space Z values, one per point - sample them out of a depth map with 'CV Sample Array At Points'. A single value broadcasts to every point, which is how you unproject onto a plane at a known distance. | |
| K | NPARRAY | 3x3 camera matrix the points were seen through. | |
| dist_coeffsopt | NPARRAY | Optional lens distortion of that camera. When wired, the points are undistorted first (cv2.undistortPoints), so pixels measured on a real photo unproject onto the ideal ray. Leave unwired for a pinhole render. |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| points_3d | NPARRAY | Nx3 float32 points in CAMERA space (X right, Y down, Z into the scene). To get them into the model's own frame, invert the pose that rendered them - 'OpenCV Pose To Matrix' -> 'cv2 invert' -> 'CV Transform Points 3D' - which is what the 'CV Pick Point On Mesh' blueprint wires up. |
| valid | NPARRAY | N uint8: 0 where the depth was missing (non-finite or <= 0) and the point came back NaN. A click that landed on the background reads 0 here. |
| count | INT | How many points unprojected successfully. |