CV Depth to 3D Points
The two lines of math, and the unit trap
- depth
- camera_matrix
- colors
- points
- colors
CV Depth to 3D Points back-projects a depth map into a coloured 3D point cloud using camera intrinsics. It's the bridge between the 2D half of your workflow (any of the depth models everyone runs) and the 3D half (a .ply you can open in MeshLab, Blender, or a viewer).
You want this when you're doing parallax or VR from a single image and want actual geometry rather than a smeared warp, when you're printing a bas-relief (a depth map already is a height field), or when you want to sanity-check a calibration. It's also the honest way to find out that your "depth map" is inverted, because the cloud will be inside out.
How it works
It is exactly the pinhole back-projection, per pixel:
X = (u - cx) * Z / fx
Y = -(v - cy) * Z / fy
Z = depth * depth_scale
Two details in the implementation are worth knowing because they explain otherwise-mysterious behaviour. First, Y is negated. Image coordinates grow downward, camera coordinates grow up, so without the flip your cloud renders upside down in a viewer and looks like the node is broken. Second, the node drops non-finite rows: a rasterized depth map is +inf where nothing was drawn, and at the principal point (u - cx) is exactly zero, so 0 * inf = NaN - that one gets masked out quietly rather than poisoning the array.
Note what this isn't: cv2.reprojectImageTo3D exists for stereo disparity, and it wants a disparity map and a Q matrix. This node handles monocular depth directly, which is the case you almost always have in ComfyUI.
Inputs and outputs
depth- HxW float32, single channel. Values can be in any unit; that's whatdepth_scaleis for. Load throughImage → CV Arrayin GRAY float32 and the channel is already extracted.camera_matrix- the 3x3 K, fromCV Calibrate Camera (Chessboard)or built withCV Camera Matrix. It readsfx = K[0,0],fy = K[1,1],cx = K[0,2],cy = K[1,2].depth_scale- the multiplier that turns your depth values into metres.0.001if the map is in mm,1.0if it's already metres,65.535for the NYU Depth V2 convention of uint16 millimetres loaded as 0–1.colors(optional) - an HxWx3 or HxW image the same size as the depth. Skip it and every point is white, which is fine for geometry and useless for anything you want to look at.
Out come points (Nx3 float32) and colors (Nx3 uint8 RGB). Wire them into CV Write PLY (Point Cloud) and you have a file. From there the core 3D nodes can preview it - though note CV Downsample Point Cloud and CV Filter Point Cloud exist because a full-res cloud from a 2MP depth map is four million points and most viewers will simply stop responding.
Install
It's part of comfyui_cv (bmad4ever/comfyui_cv). ComfyUI Manager → search "ComfyUI CV", or:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"
Restart ComfyUI. Python ≥ 3.12, V3 node API ComfyUI. The contrib wheel is mandatory, not a suggestion - a non-contrib opencv-python in the same site-packages/cv2 silently removes the contrib nodes, and tools/repair_opencv_contrib.py --check detects it.
Where people get burned
- The cloud is metres wide or 1000 metres wide. Almost always
depth_scale. Get the unit of your depth map from whoever made it; there is no way to infer it later. - The cloud looks right but its measurements are nonsense. Most community depth models are trained to output relative depth - MiDaS lineage and Depth Anything included - so a scene comes out up to an unknown scale and offset. Multiply that by a real focal length and you get a plausible-looking object at the wrong size. If you need metric, use a metric model from the ZoeDepth lineage and know that even those are approximate. There's a useful rule the depth-model comparisons keep re-teaching: any single-image depth, however pretty, is a picture of depth, not a measurement.
- Everything is squashed or stretched. A K matrix from a different resolution, or from a different camera, than the image the depth was computed for. Intrinsics scale with resolution - half the width means half the
fxandcx. - The cloud is inverted. Check the depth convention: some models emit inverse depth (near = bright) rather than depth (near = dark). No node can guess this for you; flip the input with
cv2_normalize/cv2_subtractorCV Contrast (CLAHE/Equalize).
Inputs (4)
| Name | Type | Default | Description |
|---|---|---|---|
| depth | NPARRAY | HxW float32 depth map. Values can be in any unit (mm, meters, raw model output) — use depth_scale to convert. Single-channel; if loaded via Image->CV (GRAY float32), the channel is already extracted. | |
| camera_matrix | NPARRAY | 3x3 camera intrinsic matrix K from Calibrate Camera (Chessboard) or CV Camera Matrix. fx=K[0,0], fy=K[1,1], cx=K[0,2], cy=K[1,2]. | |
| depth_scale | FLOAT | 1.000 | Multiplier to convert depth values to meters. E.g. 0.001 if depth is in mm, 1.0 if already in meters, 65535/1000=65.535 for NYU Depth V2 (uint16 mm loaded as float32 0-1). |
| colorsopt | NPARRAY | HxWx3 or HxW uint8/float RGB color image (same dimensions as depth). If disconnected, all points are white. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| points | NPARRAY | Nx3 float32 array of 3-D positions. |
| colors | NPARRAY | Nx3 uint8 array of per-vertex RGB colors. |