cv2.depthTo3d
A depth map in, a point cloud out
- depth
- K
- mask
- nparray
This is a back-projection: for every pixel with a depth value, work out where that pixel is in 3D. Give it a depth map and the camera's intrinsics and it returns an H×W×3 array of XYZ coordinates - the geometric step in front of point clouds, parallax rigs and 3D-from-a-photo tricks.
The maths is the pinhole model, and it's worth seeing because it explains every surprising result you'll get: X = (u − cx) · Z / fx, Y = (v − cy) · Z / fy, Z = depth. Two linear scales and two offsets - the focal lengths and the principal point from K - plus the depth itself. If your cloud is stretched, the focal lengths are wrong. If it's shifted, the principal point is. If it's the wrong size entirely, it's the depth units.
Inputs and output
depth accepts an NPARRAY link or a ComfyUI IMAGE/MASK directly. K is the 3×3 intrinsic matrix and accepts only an NPARRAY - the pack's tooltip is explicit that it wants a data matrix, and adds what you should expect back: "the result is an (H, W, 3) XYZ map in camera coordinates". mask is optional and restricts which pixels get converted, which is how you avoid spending time on the parts of a depth map that are holes.
One output: a single NPARRAY holding that raster XYZ map. Because it's a raster (one row per image row) rather than a flat list of points, it's not directly PLY-ready - but it drops straight into the pack's array nodes for slicing, reshaping and per-pixel math.
The units question, which is the whole ballgame
The RGB-D depth convention this comes from is a real measurement: 16-bit integers in millimetres, or float32 in metres. Hand it a normalised 0–1 AI depth map - the kind MiDaS, Depth Anything and friends produce - and you'll get a cloud that's a few units across, with everything crammed into a sliver near the camera. Relative depth from a monocular model has no metric scale at all; you have to impose one (the pack's curated depth node has a depth_scale input for exactly that, with 0.001 for millimetre data and 65535/1000 for NYU-style uint16-as-float).
The curated node you'll probably use instead
The same pack ships CV Depth to 3D Points, and for most workflows it's the better door: it takes a depth map, a camera_matrix and optional colour, and returns an N×3 float32 point array plus N×3 colours, with the NaN rows from invalid pixels already dropped and a depth_scale multiplier for unit conversion. That output is what CV Write PLY wants, so the whole "photo plus depth map becomes a mesh you can open in Blender" pipeline is three nodes.
Reach for the raw cv2.depthTo3d when you specifically want the raster map: feeding per-pixel math in the array domain, comparing a depth camera against a synthetic render, or working in stereo-adjacent territory where the alignments line up with the pack's registerDepth-style calibration nodes. One honest caveat from the wider AI-depth world: monocular depth models give you relative depth with a per-image scale, so a cloud built straight from one is great for parallax-style visual effects and useless as a measurement.
Install
Manager → Install Custom Nodes → ComfyUI CV (publisher bmad4ever), or:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"
Restart ComfyUI. Python ≥ 3.12 and a recent V3-API ComfyUI; the contrib headless wheel is the only dependency and no models are downloaded for this node. The pack is GPL-3.0, forked from geroldmeisinger/opencv-comfyui, strongly LLM-assisted, and its author's own disclaimer warns the example pipelines are overfitted to their test datasets - read any geometry node's output with that in mind.
Common issues
- The
Klink refuses to connect. It's NPARRAY-only. Build it withCV Camera Matrixor a calibration node, or parse a literal withCV Parse Matrix. - The cloud is a tiny speck. Your depth map is normalised 0–1 rather than metric. Scale it, or switch to the curated node with
depth_scale. - NaN or Inf coordinates. Zero-depth pixels - the classic hole in an AI depth map or a missing reading. Use the
maskinput to exclude them, or filter the result. - The cloud is mirrored vertically relative to the image. Camera conventions: OpenCV's
Ypoints down in image space, and various 3D viewers and export formats assumeYup.CV Flip Axisexists in this pack for that.
Inputs (3)
| Name | Type | Default | Description |
|---|---|---|---|
| depth | NPARRAY,IMAGE,MASK | - - - Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| K | NPARRAY | - - - A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here. | |
| maskopt | NPARRAY,IMAGE,MASK | - - - Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| nparray | NPARRAY | — |