Nodes/ComfyUI CV/cv2.reprojectImageTo3D
ComfyUI Node

cv2.reprojectImageTo3D

Turn disparities into an XYZ map you can actually use

By bmad4ever·Created 3 months ago·Updated 14 days ago· 1
cv2.reprojectImageTo3D
  • disparity
  • Q
  • nparray
◄handleMissingValuesfalse►
◄ddepthsame as input►

The one line that turns stereo into 3D

A disparity map says "this pixel moved 43 columns between the left and right camera". That's a measurement in image units. Multiply through the stereo geometry and it becomes a distance. cv2.reprojectImageTo3D is that multiplication, done for every pixel at once: disparity + the Q matrix → a three-channel float image where each pixel holds (X, Y, Z) in the rectified first camera's coordinate frame.

If you've ever wanted a point cloud out of a pair of photos without training a depth model, this node is the hinge. It's also the step where a stereo pipeline stops being 2D: everything before it is images, everything after it is geometry.

Inputs

  • disparity - single channel, and it accepts 8-bit unsigned, 16-bit signed, 32-bit signed or float32. Same size as the rectified image.
  • Q - the 4×4 perspective transformation from cv2.stereoRectify (or CV Stereo Rectify (Uncalibrated)'s outputs if you build it). This is where the baseline and focal length live, so a Q assembled by hand with the wrong baseline gives a cloud that's wrong in exactly one axis and looks fine otherwise.
  • handleMissingValues (optional, default false) - for the pixels where disparity was never computed. Turn it on and those pixels become 3D points at a huge Z (the docs say 10000), which is a sentinel rather than a distance: filter them out downstream. Leave it off and you get whatever NaN/Inf arithmetic produces, which flows into a cloud and then into a PLY and then into someone's Blender scene.
  • ddepth (optional, default same as input) - the output array's depth. same as input is -1 → CV_32F, which is what you want; the alternative explicit depths (CV_16S, CV_32S) exist for fixed-point pipelines and are a good way to quantise your geometry for no reason.

Output: nparray, (H, W, 3) float - the XYZ map. Not an image, even though cv2 calls it _3dImage; it's a data array with three floats per pixel, and previewing it as a picture shows coloured noise.

Making it useful

An (H, W, 3) array is not yet a point cloud. The pack's own stereo example does the honest chain, and it's worth copying:

CV Stereo Disparity → CV Disparity Interpolate → cv2.reprojectImageTo3D → CV Reshape Array → CV Filter Points 3D By Mask / CV Filter Point Cloud → CV Flip Axis (3D) → CV Write PLY → CV Preview 3D (Calibrated Camera).

The reshape flattens (H, W, 3) to N×3 - that's the shape every 3D node in the pack speaks, and it's the single most common missing step for first-timers. The filters throw out far/outlying points (and the sentinel depths from handleMissingValues). The axis flip exists because OpenCV's camera frame is Y-down, Z-into-the-scene while glTF/three.js is Y-up - the pack's CV Convert Axis Convention (3D) node documents that negating one axis alone is a reflection that nothing downstream catches. Choose your convention deliberately, once.

For a plain depth map - not a disparity - you don't need this node at all: CV Depth to 3D Points back-projects a depth image, and cv2.depthTo3d is the wrapper if you want the array-level version.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"

Manager: search "ComfyUI CV" (bmad4ever). Python ≥ 3.12, and a ComfyUI recent enough for the V3 node API - the pack's nodes register through comfy_api.latest and won't load on older builds. Restart ComfyUI, reload the page.

Troubleshooting

The cloud is a flat plane, or a wedge. Your Q doesn't match the disparity. The classic version of this is scale: some stereo matchers emit disparity multiplied by 16 (fixed point), and if that scaling isn't undone before this node, every Z is 16× off, which looks like a "compressed" scene rather than an error. Feed disparity in the units Q expects - the pack's examples put CV Disparity Interpolate between the matcher and here for exactly this kind of plumbing.

A few points at absurd distances. That's handleMissingValues doing what it says. Filter by distance (CV Filter Point Cloud) or by the valid mask, and don't let Z=10000 into an export.

Everything is NaN. Rectification didn't happen (the disparity is not from a rectified pair) or Q is wrong. reprojectImageTo3D assumes rectified coordinates; if you're matching unrectified images, the epipolar lines aren't horizontal and no amount of Q fixes it. There's also cv2.patchNaNs and cv2.finiteMask in the pack if you need to clean a partially-NaN cloud.

Preview looks like TV static. It's a float XYZ array, not an RGB image. Use CV Reshape Array + a viewer, or the calibrated 3D preview node.

Categoryimage/CV/low-level/cv2 R

Inputs (4)

NameTypeDefaultDescription
disparityNPARRAY,IMAGE,MASKInput single-channel 8-bit unsigned, 16-bit signed, 32-bit signed or 32-bit floating-point disparity image. The values of 8-bit / 16-bit signed formats are assumed to have no fractional bits. If the disparity is 16-bit signed format, as computed by or Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
QNPARRAY$4 \times 4$ perspective transformation matrix that can be obtained with A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here.
handleMissingValuesoptBOOLEANfalseIndicates, whether the function should handle missing values (i.e. points where the disparity was not computed). If handleMissingValues=true, then pixels with the minimal disparity that corresponds to the outliers (see StereoMatcher::compute ) are transformed to 3D points with a very large Z value (currently set to 10000). Preset to the OpenCV default (False).
ddepthoptCOMBOsame as inputThe optional output array depth. If it is -1, the output image will have CV_32F depth. ddepth can also be set to CV_16S, CV_32S or CV_32F. The function transforms a single-channel disparity map to a 3-channel image representing a 3D surface. That is, for each pixel (x,y) and the corresponding disparity d=disparity(x,y) , it computes: $$\begin{bmatrix} X \ Y \ Z \ W \end{bmatrix} = Q \begin{bmatrix} x \ y \ \texttt{disparity} (x,y) \ 1 \end{bmatrix}.$$

Outputs (1)

NameTypeDescription
nparrayNPARRAY—