Nodes/ComfyUI CV/CV Decompose Homography
ComfyUI Node

CV Decompose Homography

Getting camera motion out of a homography

By bmad4ever·Created 3 months ago·Updated 14 days ago· 1
CV Decompose Homography
  • homography
  • camera_matrix
  • ref_points
  • rotations
  • translations
  • normals
  • visible
  • count
  • best_index
  • found

What it's for

Two photos of the same flat thing - a wall, a poster, a floor, a chessboard - and CV Find Homography gives you the 3×3 that maps one onto the other. Fine. But the homography is a projection of motion, and if you're placing something in a scene, stabilising a shot, or reconstructing where the camera went, what you actually want is the motion: rotation, translation, and where that plane is pointing.

That's what decomposing the homography gets you, and it's the step most people skip because the answer is annoying. Given H and the camera's intrinsic matrix K, the motion is recoverable - up to four candidate solutions, because a homography between two views of a plane is genuinely ambiguous. Two of the four are mirror pairs a real camera can't produce, and killing them is what the rest of this node is for.

How it works

Under the hood it's cv2.decomposeHomographyMat, plus filterHomographyDecompByVisibleRefpoints when you give it the points to filter with. The logic: a physically valid solution has the plane in front of both cameras. A candidate that puts the plane behind one of them is impossible, so it can be discarded - and since the ambiguity comes in mirror pairs, discarding the impossible ones normally leaves you with exactly two.

Two survivors is not a bug. It's the real two-fold planar ambiguity: two different (R, t, n) combinations produce the same image of the same plane, and no filtering resolves it from two views. You need a third view, or an outside fact about the scene.

Optional ref_points - at least four Nx2 pixel points in the source image lying on the plane, normally the very correspondences H was fitted on - trigger the filter; they're normalised with K and pushed through H internally. Without them, visible comes back all-255 and nothing has been filtered.

The outputs, and why they're shaped this way

  • rotations - (N, 3, 3) float64, up to four candidate rotations.
  • translations - (N, 3), and read the units carefully: these are scaled by the plane distance, i.e. t/d, not t. Only the direction is metric unless you independently know how far away the plane is. Multiply by the real distance and you get units.
  • normals - (N, 3) unit-length plane normals, in the source camera frame.
  • visible - (N, 1) uint8, 255 for each candidate that keeps the reference points in front of both cameras. Expect two.
  • count - usually 4, or 0 if nothing was found.
  • best_index - the index of the first surviving candidate; 0 when no ref_points were supplied, -1 when nothing was found or everything was rejected. It is not necessarily the right answer.
  • found - a boolean, and the thing you branch on.

Pulling one candidate out is CV Index Batch with best_index wired in (widget-convert its index input to get the socket). Then, to actually choose between the two survivors, use an outside cue: dot the candidate normals against a plane normal you already believe in, and let CV Pick Value (sorted) pick the winner.

The repo's 56_matrix_decomposition.json builds the whole thing.

Install

ComfyUI Manager, search ComfyUI CV, or:

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv

Restart afterwards. Python ≥ 3.12, a ComfyUI recent enough for the V3 node API, and:

pip install "opencv-contrib-python-headless~=5.0.0.93"

Where people get burned

Using best_index as the answer. It's the first survivor, not the correct one. Check visible for the other, and break the tie yourself - this is the node's own advice and it's the difference between a working AR placement and one that's subtly inside out.

Reading translations as metres. They're divided by the plane distance. If the numbers look absurdly small, that's why: they're ratios, not distances.

Branching on -1. When nothing is found, best_index is -1 and CV Index Batch clamps a -1 to 0 - so you'd silently take the first (empty) candidate. Check found first, always.

A degenerate H. Too few or collinear correspondences give you a degenerate homography, and this decomposition is one of the calls that will happily return nonsense for one. The node is failure-tolerant: found = false and empty stacks instead of an exception. Convenient - and it means nothing shouts that your matching was bad. CV Inspect Data on the H is how you notice.

ref_points from the wrong image. They must be the source-frame pixel points that sit on the plane. Feed source points from the destination image, or a mix, and the visibility filter will cheerfully reject the correct candidates.

Same K for both views. K must be the intrinsics of the camera that took both images; different zoom or different cameras and you're solving a problem you don't have.

Categoryimage/CV/features

Inputs (3)

NameTypeDefaultDescription
homographyNPARRAY3x3 homography mapping the SOURCE view to the DESTINATION view ('CV Find Homography', cv2.getPerspectiveTransform, or a composed matrix).
camera_matrixNPARRAY3x3 intrinsic matrix K of the camera that took BOTH views ('CV Camera Matrix' / 'Calibrate Camera').
ref_pointsoptNPARRAYOptional Nx2 (or Nx1x2) PIXEL points in the SOURCE image that lie on the plane - normally the very correspondences H was fitted on. At least 4 are needed. They are normalized with K and mapped through H internally, then used to reject the candidates that would put the plane behind a camera (filling 'visible' and 'best_index').

Outputs (7)

NameTypeDescription
rotationsNPARRAY(N, 3, 3) float64 stack of candidate rotation matrices, N up to 4. Use 'CV Index Batch' to take one.
translationsNPARRAY(N, 3) float64 stack of candidate translations, SCALED BY THE PLANE DISTANCE: each row is t/d, not t. Only its DIRECTION is metric unless you know d - multiply by the real distance to the plane for units.
normalsNPARRAY(N, 3) float64 stack of candidate plane normals, in the SOURCE camera frame (unit length).
visibleNPARRAY(N, 1) uint8 mask, 255 for each candidate that keeps ref_points in front of both cameras. All 255 when no ref_points were given. Expect TWO survivors out of four: the filter removes the mirror pair, not the genuine two-fold planar ambiguity.
countINTHow many candidate solutions were returned (0 when not found, otherwise usually 4).
best_indexINTIndex of the FIRST candidate that survives the visibility filter (0 when no ref_points were given, -1 when nothing was found or every candidate was rejected). It is not necessarily the right one - check 'visible' for the other survivor, and pick between them with an outside cue (e.g. the dot product of 'normals' with a known plane normal, then 'CV Pick Value (sorted)'). 'CV Index Batch' clamps a -1 to 0, so branch on 'found' first.
foundBOOLEAN—