Nodes/ComfyUI CV/CV PPF Pose Estimation
ComfyUI Node

CV PPF Pose Estimation

Find a 3D model in a scene with no initial guess

By bmad4ever·Created 3 months ago·Updated 14 days ago· 1
CV PPF Pose Estimation
  • model
  • scene
  • pose
  • poses
  • votes
  • pose_count
  • found
◄relative_sampling_step0.050►
◄relative_distance_step0.050►
◄num_angles30►
◄relative_scene_sample_step0.20►
◄relative_scene_distance0.050►
◄max_poses5►
◄normal_neighbors12►

This is the heavy one in the cloud-matching corner of the pack. You have a model - a mesh, a scanned part - and a scene cloud that may contain it somewhere, at an unknown pose, among clutter. PPF finds it without being told where to start.

How it works

Point Pair Features: the detector hashes oriented pairs of points from the model, quantising the distance and angles between them into a lookup table. Then the scene's own point pairs vote for matching model pairs, and every vote implies a candidate 4x4 pose mapping model → scene. Votes accumulate into clusters, and the highest-voted clusters come out.

That's the crucial difference from CV ICP Register, which refines a pose you already roughly know. PPF needs no initial guess - and pays for it by being coarse. On the pack's own example it lands about 14 degrees off, which is a placement, not a fit. The standard pipeline is therefore PPF then ICP: take pose, transform the model with CV Transform Points 3D, recompute normals with CV Point Cloud Normals, run ICP against the scene, and compose the two transforms with CV Matrix Multiply.

The thing that will actually cost you time

Both clouds must be Nx6 - positions and normals. The node accepts Nx3 and computes normals itself, but don't let it: hand cv2 a bare cloud and it runs 31x to 100x slower and returns a silently wrong pose. Put CV Point Cloud Normals in front of both inputs, explicitly. If your cloud came from a mesh, use CV Mesh Vertex Normals instead - triangle winding gives normals a real sign, and a plane fit can't orient a closed object, which inverts the PPF result by roughly 180 degrees.

Inputs and outputs

model is the thing you're looking for, scene is where you're looking. Then the knobs, in order of how much they'll hurt you:

  • relative_sampling_step (0.05) - the cost dial. It's the model downsample distance as a fraction of the model's diameter, and halving it roughly quadruples training time: 3.5 seconds at 0.05 becomes about 26 seconds at 0.025 on a 2400-point model. Cost grows with the square of the sampled point count, so downsample big clouds first.
  • relative_distance_step, num_angles, relative_scene_sample_step, relative_scene_distance - the hash-table quantisation and how thoroughly the scene is searched. Coarser is faster and blurrier; relative_scene_sample_step 0.2 means every fifth scene point is used as a reference.
  • max_poses (5) - how many candidates come back.

normal_neighbors only matters if a cloud arrives without normals, which per the section above is the case you want to avoid.

Outputs: pose (the 4x4 for the top cluster, identity when found is false), poses (a Kx4x4 stack of the top candidates), votes (a Kx1 vote count - a relative confidence that scales with cloud size and sampling steps, not a quality score), pose_count, and found. Take the ranked list seriously: PPF often has the right answer at rank two or three, so refine several candidates with ICP and keep whichever gives the lowest CV Point Cloud Nearest Distance. Gate everything on found with an if/else node - a degenerate cloud (a plane, too few points) returns identity rather than an exception.

Installing

Pack-wide. ComfyUI Manager → search "ComfyUI CV", or:

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
# restart ComfyUI

Python ≥ 3.12, a recent V3-API ComfyUI, and opencv-contrib-python-headless~=5.0.0.93. Contrib is non-negotiable here: ppf_match_3d only exists in the contrib build, and installing a plain opencv-python over the contrib wheel silently strips it. Run the pack's checker if nodes go missing:

python tools/repair_opencv_contrib.py --check

Where it goes wrong

Slow, and it's always the sampling step. Big cloud, default steps, and you're waiting minutes - read the timing note above before you blame the pack.

Wrong pose with high votes. Votes are not confidence. If PPF is confidently wrong, check the normals' sign first, then whether the model and scene are in the same coordinate convention (a mirrored axis convention gives a mirrored pose).

And remember this pack is explicitly not production-ready: the author's own README warns the workflows are showcase-grade and the code was written with heavy AI assistance. Treat PPF here as a genuinely useful OpenCV wrapper with a coarse answer, and verify the pose before you build on it.

Categoryimage/CV/contrib

Inputs (9)

NameTypeDefaultDescription
modelNPARRAYThe object to look for, Nx6 (x,y,z,nx,ny,nz) from 'CV Point Cloud Normals'. An Nx3 cloud is accepted and its normals are computed here - never hand cv2 one directly, it silently runs 31x-100x slower and answers wrong.
sceneNPARRAYThe cloud to search in, Nx6 (or Nx3, as above). May contain clutter and other objects - that is what PPF is for.
relative_sampling_stepFLOAT0.0500.005–0.5Model downsampling, as a fraction of the model's diameter: 0.05 keeps points ~5% of the diameter apart. THE cost knob - halving it roughly quadruples training time (0.025 -> 26 s where 0.05 -> 3.5 s on a 2400-point model).
relative_distance_stepFLOAT0.0500.005–0.5Quantization of the point-pair distance in the hash table, as a fraction of the diameter. Larger tolerates more noise and blurs the vote.
num_anglesINT306–120Angle bins the pair features are quantized into. More = finer rotation resolution, bigger table.
relative_scene_sample_stepFLOAT0.200.01–1Fraction of scene points used as reference points (0.2 = every 5th). Lower = faster, less thorough.
relative_scene_distanceFLOAT0.0500.005–0.5Scene downsampling distance, as a fraction of the model diameter (the scene equivalent of relative_sampling_step).
max_posesINT51–100How many of the highest-voted pose clusters to return in 'poses'. 'pose' is always the first.
normal_neighborsoptINT123–200Neighbours per plane fit, used ONLY for an input that arrives without normals. Prefer wiring 'CV Point Cloud Normals' explicitly.

Outputs (5)

NameTypeDescription
poseNPARRAY4x4 rigid transform mapping model -> scene, the highest-voted cluster (identity when found=false). Feed 'CV Transform Points 3D'.
posesNPARRAYKx4x4 stack of the top 'max_poses' candidates, best first. PPF often gets the right pose at rank 2-3, so refine several with ICP and keep the one with the lowest 'CV Point Cloud Nearest Distance'.
votesNPARRAYKx1 float32 vote count per candidate - a RELATIVE confidence only (it scales with the cloud size and the sampling steps), not a quality measure.
pose_countINTNumber of candidates returned (<= max_poses).
foundBOOLEANFalse when the match could not run (empty or degenerate cloud, cv2 error). Gate the transform on it with a 'Basic data handling: IfElse'.