Nodes/ComfyUI CV/cv2.estimateTranslation3D
ComfyUI Node

cv2.estimateTranslation3D

When you know the clouds only shifted

By bmad4ever·Created 4 months ago·Updated 16 days ago· 1
cv2.estimateTranslation3D
  • src
  • dst
  • retval
  • out
  • inliers
◄ransacThreshold3.0000►
◄confidence0.9900►

Three degrees of freedom instead of twelve. estimateTranslation3D fits a pure translation between two corresponding 3D point sets - no rotation, no scale, no shear - and if that sounds too restrictive to be useful, that's exactly why it's here. The constraint is the method. When you know the answer is a shift, refusing to model rotation means the fit can't soak up noise as a phantom twist, and it needs corresponding points that are far easier to obtain than the full 3D spread a general affine registration demands.

The mechanism

Given N x 3 pairs, it solves for the single 3-vector that best maps src onto dst, under RANSAC with the same consensus machinery as its affine sibling: sample minimal subsets, count points within ransacThreshold, keep the largest consensus. The minimum sample is tiny - a translation is one unknown vector - which is why this is the most forgiving registration in the family. A handful of well-spread correspondences will do it.

The price is obvious: if the two sets are actually rotated relative to each other, the residual is baked into the translation you get back, and the inlier count will tell you (badly). Treat it as a hypothesis test as much as a fit: "is the difference between these clouds a shift?" If half your points come back as outliers, the answer is no, and you want cv2.estimateAffine3D or the curated CV Register Point Clouds (3D) instead.

Inputs and outputs

src and dst are NPARRAY only, N x 3, same length, corresponding point-for-point. dst carries the family's in-place warning: the underlying cv2 call writes into it, but this wrapper passes cv2 a private copy, so your input array is not modified - a detail worth trusting, since the C++ signature looks like it would clobber your data.

The optional knobs are ransacThreshold (default 3.0) and confidence (0.99). Same caveat as everywhere in this family: the threshold is in the cloud's own units, so 3.0 is a sane pixel-ish default and a nonsense metre default. Set it to something like your depth noise floor - centimetres for a good stereo rig, tenths of a metre for a noisy one.

Outputs are retval (a BOOLEAN), out (the translation), and inliers (the N×1 mask). out is a 3-vector rather than a matrix, so it doesn't drop straight into CV Transform Points 3D; build the 3×4 or 4×4 around it if you need to apply it, or keep it as a measurement - the displacement between two clouds is a number you might want to read and act on rather than a transform you want to apply.

Where it fits

Rig drift, temporal offset, and calibration checks. Two clouds of the same static scene that should be identical but aren't; a scan that has slipped relative to a reference between passes; the sanity check on a rig whose baseline you wrote down and want verified. In this pack's 3D neighborhood it pairs with CV Depth to 3D Points, CV Merge Point Clouds, and CV Transform Points 3D, and the constraint math (how much of an apparent motion is translation versus rotation) is the sort of thing that shows up whenever depth clouds are compared - depth-estimation.md is background on how noisy those inputs are in the first place.

If you're new to the family: read this one as the "cheapest hypothesis" version, cv2.estimateAffine3D as the general one, and the curated CV Register Point Clouds (3D) as the one with the plumbing and the failure-tolerant found flag that most graphs should actually use.

Installing it

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"

ComfyUI Manager: search comfyui_cv, install, restart. Python ≥ 3.12, V3-API ComfyUI. Node path: image/CV/low-level/cv2 E.

Traps

Points that are all clustered in one small region will fit a translation fine and be wrong about the world - spread matters, even for three unknowns. Zero or one correspondence should be treated as "no answer", and the node may fail rather than return something useless. And if inliers comes back mostly zero while out looks plausible, believe the mask: a translation fit on a rotated pair of clouds is the most plausible-looking wrong answer in this entire family.

Categoryimage/CV/low-level/cv2 E

Inputs (4)

NameTypeDefaultDescription
srcNPARRAY - - - A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here.
dstNPARRAY - - - The low-level cv2 function writes its result into this array in place, but this wrapper passes cv2 a private copy, so your input array is never modified. A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here.
ransacThresholdoptFLOAT3.0000-1e+38–1e+38 - - - Preset to the OpenCV default (3.0).
confidenceoptFLOAT0.9900-1e+38–1e+38 - - - Preset to the OpenCV default (0.99).

Outputs (3)

NameTypeDescription
retvalBOOLEAN—
outNPARRAY—
inliersNPARRAY—