Nodes/ComfyUI CV/CV Feature Extract (Model)
ComfyUI Node

CV Feature Extract (Model)

DISK and SuperPoint keypoints inside ComfyUI

By bmad4ever·Created 3 months ago·Updated 14 days ago· 1
CV Feature Extract (Model)
  • image
  • keypoints
  • descriptors
  • count
◄model►
◄max_keypoints2048►
◄resize_long_side1024►
◄engineauto (default engine)►

Classic detectors - SIFT, ORB, AKAZE - are fine until the lighting changes, the viewpoint swings, or the texture is soft. Learned extractors do noticeably better on exactly those cases, and DISK and SuperPoint are the two that the OpenCV-adjacent ecosystem ships as ready ONNX exports. This node runs one on your image and returns keypoints plus descriptors in the standard format the pack's matcher and drawing nodes expect.

You don't have to use it for matching photos. It's the front end of any "find the same object across two images" pipeline, which in ComfyUI usually means: find a reference object, get its scale and rotation, and then warp, mask, or composite based on that.

How it works

This is cv2.dnn running an ONNX model, in the pack's interruptible DNN worker (feature_extract task) so a slow run can be cancelled. Two implementation points are worth your attention because they leak into the results:

Preprocessing follows the LightGlue reference implementation, read off the model rather than assumed. Pixels are scaled to 0..1, fed as RGB or gray according to the channel count the model declares, and the image is padded up to a multiple of 16 - necessary because DISK is a U-Net and its downsampling doesn't tolerate arbitrary sizes. If you're comparing outputs against a reference implementation elsewhere, this is why they agree.

Output parsing handles several shapes, because ONNX extractors aren't standardised: three-output models (DISK's keypoints (1,N,2) + scores (1,N) + descriptors (1,N,D)), flat two-output models ((1,N,4) + (1,N,D)), and dense heatmap-style models ((1,C,H,W) + (1,C,H,W)). Keypoints always come back in the input image's pixel coordinates, even though the network saw a resized, padded version.

Zero features is a valid result - empty outputs, not an error.

Inputs and outputs

  • image - BGR or gray.
  • model - a combo listing every .onnx in ComfyUI/models/onnx, subdirectories included; there's no enforced path convention, so you pick the right extractor for your matcher. The convention in the pack's own example is lightglue/disk.onnx and lightglue/superpoint.onnx.
  • max_keypoints (default 2048, 64–16384) - keeps the highest-scoring ones. The models emit keypoints in detection order, not by score, so this is doing real work, not trimming a tail.
  • resize_long_side (default 1024, 0 = native) - the LightGlue reference resizes so the long side is this long, and it upscales small images too. That's where most of the keypoints on a small photo come from, so don't set it to 0 reflexively.
  • engine (optional) - the DNN engine for inference, default auto.

Outputs are keypoints (CV_KEYPOINTS), descriptors (an (N, D) float32 array) and count. The keypoints go into the pack's keypoint pipeline - CV Filter Keypoints, CV Match Features (Model), CV Draw Keypoints, CV Draw Matches. Note the pairing rule: keypoints and descriptors must stay together, so filter them with CV Filter Keypoints rather than splitting the array yourself.

Getting models

They aren't bundled. The pack's model_sources.txt points at the LightGlue-ONNX releases (fabio-sim's repo, v0.1.0) for disk.onnx, superpoint-family extractors and the disk_lightglue.onnx matcher that pairs with DISK; licences there are Apache-2.0. Drop them under:

ComfyUI/models/onnx/lightglue/

Then reload the page or restart ComfyUI - the model combo is built from what's on disk when the node definitions load, so a file copied mid-session doesn't appear until you refresh.

Install

# ComfyUI Manager → search "ComfyUI CV" → install → restart
# or manually:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
cd comfyui_cv && pip install -r requirements.txt

Python ≥ 3.12 and a recent ComfyUI built on the V3 node API. The only dependency is opencv-contrib-python-headless~=5.0.0.93 (with numpy and torch); keep it the contrib wheel, since installing plain opencv-python over it silently strips the contrib submodules and removes contrib nodes.

What goes wrong

  • Empty model dropdown. No .onnx in ComfyUI/models/onnx (subdirectories count), or you didn't refresh the page.
  • The model loads but produces garbage. You've paired an extractor with the wrong matcher, or fed a matcher model into the extractor slot. DISK descriptors belong with a DISK-family matcher.
  • A few hundred features on a small image. resize_long_side is doing its job; if you actually want native-resolution behaviour, set it to 0 and expect fewer, larger-scale features.
  • Model won't load at all. The README's standing caveat applies to this whole DNN family: a perfectly valid ONNX export can still be unloadable here, limited by both OpenCV's DNN implementation and the pinned OpenCV version. That's not a path typo.
  • Don't reach for this if classic features work. CV Detect Features running SIFT/ORB costs nothing to download and is often enough; come here when the matches are the problem, not before.

The pack is bmad4ever's fork and rewrite of geroldmeisinger's opencv-comfyui, written - by the author's own very explicit disclaimer - with heavy LLM assistance and not recommended for production without your own review. This node concentrates that risk more than most, since the parsing heuristics are the fragile part. Test on your own images.

Categoryimage/CV/dnn

Inputs (5)

NameTypeDefaultDescription
imageNPARRAY,IMAGEImage to extract features from (BGR or gray). Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
modelCOMBOONNX feature extractor model from models/onnx (e.g. lightglue/disk.onnx, lightglue/superpoint.onnx).
max_keypointsINT204864–16384Maximum keypoints to return, keeping the highest-scoring ones (the models emit keypoints in detection order, not by score).
resize_long_sideINT10240–8192Resize the image so its LONG side is this many pixels before extraction, the way the LightGlue reference implementation does (its own default is 1024). It UPSCALES small images too, which is where most of the keypoints on a small photo come from. 0 keeps the native resolution. Keypoints are mapped back to the input image either way.
engineoptCOMBOauto (default engine)DNN engine for inference.

Outputs (3)

NameTypeDescription
keypointsCV_KEYPOINTScv2.KeyPoint list.
descriptorsNPARRAY(N, D) float32 descriptors.
countINTNumber of detected features.