Nodes/ComfyUI CV/CV Detect Features
ComfyUI Node

CV Detect Features

ORB, SIFT, AKAZE or BRISK — detect and describe in one node

By bmad4ever·Created 3 months ago·Updated 14 days ago· 1
CV Detect Features
  • image
  • mask
  • keypoints
  • descriptors
  • count
◄detectorORB►
◄max_features2000►
◄clahefalse►

CV Detect Features is the front half of the classic matching pipeline: it finds local features and computes their descriptors in one pass, so the output can go straight into CV Match Features and then into CV Find Homography (RANSAC).

Reach for it when you need to know where two images correspond rather than just that they're similar. Alignment of two exposures, stitching, template location across a scale change, motion between frames of a video, pulling a planar object out of a scene - that's the family of jobs. The pack's panorama stitching and LightGlue-style examples all start here or somewhere very like it.

Which detector, honestly

  • ORB - the default, and the right default. Fast, free, rotation-aware, binary descriptors matched with a Hamming distance. If you don't know why you'd pick something else, pick this.
  • SIFT - slower, and still the most robust to large scale and rotation changes. Worth it when ORB's matches fall apart on a hard pair.
  • AKAZE - a decent middle ground on honestly very hard images; note it ignores max_features, as does BRISK.

The inputs and outputs

Required: image (BGR or gray, converted internally), detector (ORB / SIFT / AKAZE / BRISK), max_features (default 2000), and clahe (off by default). Optional mask restricts detection to where mask > 0.

Two things about max_features catch people. The tooltip spells out the semantics and they differ per detector: ORB with 0 means keep 0 features (empty output), while SIFT with 0 means unlimited (keep everything). AKAZE and BRISK ignore the field entirely. So "I set it to 0 to mean no limit and got nothing" is an ORB thing, and it's working as designed.

clahe is the underrated switch: on dim, low-contrast or unevenly-lit images, contrast equalisation before detection can be the difference between 40 features and 900. It's off by default because on clean images it mostly adds noise-driven keypoints.

Outputs: keypoints (a CV_KEYPOINTS list), descriptors (the NPARRAY you feed CV Match Features), and count.

How it works

The detector builds the appropriate OpenCV object - ORB, SIFT, AKAZE or BRISK - and calls detectAndCompute. That pairing is the point: descriptors are computed in the same pass as detection, in the detector's own descriptor space. ORB's are binary strings compared by Hamming distance; SIFT's are 128-float vectors compared by L2. You can't mix the two, which is why matching is done by the node downstream rather than by you.

Descriptors and keypoints travel as separate sockets precisely so you can filter one and keep them aligned - CV Filter Keypoints exists for that, and it keeps the descriptor rows in sync.

Install

From comfyui_cv (bmad4ever/comfyui_cv). ComfyUI Manager → search "ComfyUI CV", or:

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"

Restart when it's done. Python ≥ 3.12 and a recent V3-API ComfyUI. Keep the OpenCV install at contrib and keep it at the pinned 5.0.0.93 - the pack's behaviour is curated against that build, and SIFT specifically lives in the contrib set, so a plain opencv-python wheel leaves you wondering why one detector is missing. tools/repair_opencv_contrib.py --check tells you whether your site-packages/cv2 is the contrib build.

Common issues

  • Zero keypoints on a low-texture image. A smooth wall, a sky, a blurred shot. No detector finds features that aren't there. Enable clahe, lower the blur, or accept that this pair cannot be aligned - a false homography from four spurious matches is worse than an honest failure.
  • Hundreds of matches, garbage homography. Two causes: repeated texture (a brick wall, a tiled floor), where every brick matches every other brick; and outliers, which is what CV Find Homography (RANSAC) should be killing with RANSAC. Check how many inliers you actually kept.
  • Everything is mirrored or rotated in the match. Stitching generally, but worth knowing: CV Match Features and friends assume you know your own geometry. A flipped input will "match" beautifully and produce nonsense.
  • count is large but descriptors looks tiny. Some detectors simply didn't compute descriptors for every keypoint. Don't assume the two arrays have the same length without checking; the node's count refers to keypoints.

If your goal is a dense correspondence field rather than sparse points - warping one frame onto another, interpolating motion - you want optical flow (CV Optical Flow (Farneback), or the TV-L1 / DIS variants) instead. Sparse features are for finding the transform; dense flow is for describing the whole motion field.

Categoryimage/CV/features

Inputs (5)

NameTypeDefaultDescription
imageNPARRAY,IMAGEImage to detect features in (BGR or gray). Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
detectorCOMBOORBORB is fast and free; SIFT is slower but more robust to scale/rotation changes. AKAZE/BRISK ignore max_features.
max_featuresINT20000–100000Maximum keypoints to keep (ORB/SIFT only). ORB: 0 means keep 0 features (empty output); SIFT: 0 means unlimited (keep all detected features). AKAZE/BRISK ignore this input.
claheBOOLEANfalseApply CLAHE contrast equalization before detection - helps on dim or low-contrast images.
maskoptNPARRAY,MASKOptional mask: detect only where mask > 0. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.

Outputs (3)

NameTypeDescription
keypointsCV_KEYPOINTS—
descriptorsNPARRAY—
countINT—