cv2.phaseCorrelate
Sub-pixel shift between two frames, in one FFT
- src1
- src2
- shift
- response
Two images in, one number pair out: how far the second has moved relative to the first, sub-pixel accurate. No keypoints, no descriptors, no RANSAC, no convergence to babysit. Phase correlation compares the two spectra and finds the peak - which is why it works on the frames that break feature matchers: sky, water, blur, flat scans, low-texture walls.
This is the standard approach for video stabilisation, burst alignment and scan-line registration, and for ComfyUI purposes it's how you answer "how much did this frame move" without a model. It's not a warp node - it measures, you decide what to do with the measurement.
The two outputs
shift is a CV_TUPLE - composite (x, y) coordinates, sub-pixel, and the pack treats it as one value, not two. Wire it straight into any point input in the pack, or split it with CV Split Tuple for the separate numbers. (There's also CV Split Point, which is the friendlier version for pixel coordinates, with an int/float mode.)
response is a FLOAT: "Signal power within the 5x5 centroid around the peak, between 0 and 1". That's your confidence. A high response means the two frames genuinely differ by a translation; a low one means the estimator found something and you shouldn't believe it. Gate on it - the pack's curated CV Phase Correlate (Translation) exposes the same number as response and a found boolean for exactly this reason.
The two inputs, and what they must be
src1 and src2 are both documented as "Source floating point array (CV_32FC1 or CV_64FC1)". Read that literally:
- Single channel. Not colour. Convert first.
- Float. Not 8-bit. The pack's own tooltip for the iterative variant spells it out: "Single-channel FLOAT array (CV_32F or CV_64F) - not an 8-bit image."
But a ComfyUI IMAGE handed to this pack resolves to an 8-bit BGR array. So the naive wiring - Load Image on both sides, into this node - fails in cv2, not in the pack. The route that works is three nodes per side: Image → CV Array → cv2.cvtColor (to gray) → CV Cast Array (float32). Or skip all of it and use the curated node, which does gray, float and windowing internally and gives you dx, dy, response and found.
The parameter the wrapper doesn't have
cv2's phaseCorrelate takes an optional window - a Hann window to multiply both images by before correlating. The Python signature has it, the pack's own parameter documentation for this function has a tooltip for it, and the auto-generated wrapper's registry entry lists only src1 and src2. It isn't exposed.
That matters more than it sounds. Without a window, the image borders are a huge artificial edge and they can dominate the real peak - the curated node's docs say the same thing a lot more bluntly: "A Hann window is applied internally (rectangular edges alone produce a strong false peak)."
So the honest review of this raw wrapper is: it's the un-windowed estimate, useful for seeing what the primitive does, and the curated CV Phase Correlate (Translation) node is the one to put in a real workflow. If you want to window it yourself, cv2.createHanningWindow is not in the registry either (it's a create* factory, and the pack's generator only wraps top-level functions), so you'd be building the window another way - another reason to take the curated path.
Practical behaviour
Same size only. Phase correlation compares spectra; frames of different sizes have no correspondence. Crop or resize to match exactly, don't letterbox half of one.
Motion is a translation, and only a translation. Rotation, scale, perspective change or a moving subject all degrade the answer in ways the response value partly catches and partly doesn't. If the motion is more than a shift, use CV Find Transform (ECC) per the curated node's own recommendation - it's slower, but it solves for a real warp.
Uniform brightness changes don't matter. That's the one nice property here: because the comparison happens in the frequency domain after normalisation, a frame that got brighter or darker doesn't throw the measurement. Useful when you're aligning exposures rather than frames.
No batch path. Two inputs, one answer, frame 0 if you hand it a batch. For a clip you loop pairs explicitly - CV Index Batch is the pairing node - and accumulate the shifts yourself. This is the kind of thing the pack's own visual-odometry node exists to avoid rebuilding, so check the curated nodes before you commit to a loop.
Installing
ComfyUI Manager → search ComfyUI CV → Install → restart:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"
Python ≥ 3.12, recent ComfyUI (V3 node API), no models. phaseCorrelate is core cv2, so it works on any wheel.
One note on the pack itself: it's a single-author, heavily AI-assisted 2026 fork of opencv-comfyui, and I could find no community discussion of it at all - no threads, no reports, no war stories. Its README is unusually candid about the trade-offs (it explicitly warns that the auto-generated wrappers are uncurated), which is the right way to read everything here: the primitives are OpenCV's, well-tested by other people; the curated layer is the author's, and you're the first reviewer.
Inputs (2)
| Name | Type | Default | Description |
|---|---|---|---|
| src1 | NPARRAY,IMAGE,MASK | Source floating point array (CV_32FC1 or CV_64FC1) Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| src2 | NPARRAY,IMAGE,MASK | Source floating point array (CV_32FC1 or CV_64FC1) Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| shift | CV_TUPLE | - - - Subpixel coordinates (x, y) as ONE composite value - wire it straight into any cv2 point input, or into 'CV Split Tuple' for the separate numbers. |
| response | FLOAT | — |