Nodes/ComfyUI CV/CV HOG Features
ComfyUI Node

CV HOG Features

81 numbers per image, and a classifier you can actually train

By bmad4ever·Created 4 months ago·Updated 15 days ago· 1
CV HOG Features
  • image
  • features
  • dims
  • count
◄win_size20►
◄cell_size10►
◄block_size10►
◄block_stride5►
◄nbins9►

HOG is the pre-deep-learning workhorse for shape recognition: throw away colour, look at which way the edges point and how strong they are, tally that into local histograms, and you get a fixed-length descriptor of the shape rather than the pixels. It's what powered pedestrian detectors for a decade, and OpenCV's digit-recognition tutorial still ships it. CV HOG Features runs it over a batch and hands you one clean row of floats per frame.

The descriptor here is 81 numbers long with the defaults - and to be fair about the framing: 81 numbers is not going to recognise your cat. Where this node earns its keep is small, controlled, hand-made problems where you can generate your own labelled data: glyph or digit classification, simple shape sorting, distinguishing icon states, judging whether a mask is round or square. There, a HOG row plus a small cv2.ml classifier beats wiring in an ONNX model, needs no download, and trains in a second.

How it works, concretely

Every frame is converted to grayscale (three-channel input is rejected by cv2.moments-style code, so this node handles it for you), resized to win_size × win_size, and pushed through cv2.HOGDescriptor. The descriptor length isn't hardcoded - the node asks the descriptor object (getDescriptorSize()) and returns that as dims, so you can confirm rather than assume. An empty IMAGE batch comes back as an empty (0, D) array instead of an exception, which is the right call for batch pipelines.

The five geometry inputs are all author-described, and three of them are arithmetic constraints rather than taste:

  • win_size (default 20) - the square the frame is resized into. Bigger keeps more detail and grows the descriptor.
  • cell_size (default 10) - histogram cell in px. win_size and block_size must be multiples of it.
  • block_size (default 10) - normalisation block, a multiple of cell_size and at most win_size.
  • block_stride (default 5) - block step. (win_size − block_size) must be a multiple of it.
  • nbins (default 9) - orientation bins. 9 is the standard, because 180°/9 = 20° per bin is the classic choice.

Break either arithmetic rule and the node raises a configuration error instead of silently producing a garbage descriptor, which is a courtesy.

Outputs: features - (N, D) float32, one row per frame, ready for CV Train Classifier or CV Stack Feature Classes - plus dims (D, 81 with the defaults) and count (frames described).

Install

ComfyUI Manager, search the pack title comfyui_cv (bmad4ever/comfyui_cv), or:

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv

Restart ComfyUI. Two non-negotiables: Python ≥ 3.12, and a ComfyUI recent enough for the V3 node API - the pack is written entirely as io.ComfyNode classes with io.Schema schemas and defines no NODE_CLASS_MAPPINGS, so old builds simply won't load it. The one dependency:

pip install "opencv-contrib-python-headless~=5.0.0.93"

Behaviour is curated against that version. And keep it contrib - all four OpenCV wheels share one site-packages/cv2, so a non-contrib install over a contrib one silently wipes the contrib submodules and makes contrib nodes disappear from the menu. tools/repair_opencv_contrib.py --check / --apply sorts it out.

What people get wrong

A geometry error on the first run. The two multiple-of rules above, and they're strict. If you resize win_size to something awkward like 32, cell_size 10 no longer divides it.

Expecting the descriptor to be resolution-rich. win_size 20 means every image is being described at 20×20. Bumping it grows the feature vector fast - that's fine for a tiny classifier and painful for a big one.

Feeding in-the-wild photos and expecting modern accuracy. Don't. If you already have a trained network, CV Deep Features (a frozen ONNX backbone) is the honest choice. HOG is the tool for the case where you can define and generate the classes yourself.

The package-wide warning applies double to the ML corner of this pack: the README states plainly that it was written with heavy LLM assistance, that some pipelines were tuned test-first against sample data, and that it's not production-ready without your own validation. Train a classifier on this, then validate it on images it has never seen - that part is on you.

Categoryimage/CV/ml

Inputs (6)

NameTypeDefaultDescription
imageNPARRAY,IMAGEFrames to describe: an IMAGE batch gives one feature row per frame; an NPARRAY (gray/BGR/BGRA) is a single frame. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
win_sizeINT208–512Square window size in px - every frame is resized to this before computing. Bigger keeps more detail but grows the descriptor.
cell_sizeINT102–256Histogram cell size in px. win_size and block_size must be multiples of it.
block_sizeINT102–256Normalization block size in px (a multiple of cell_size, at most win_size).
block_strideINT51–256Block step in px; (win_size - block_size) must be a multiple of it.
nbinsINT92–32Orientation bins per histogram (9 is the standard).

Outputs (3)

NameTypeDescription
featuresNPARRAY(N, D) float32, one HOG row per frame - feeds 'CV Train Classifier' or 'CV Stack Feature Classes'.
dimsINTDescriptor length D (81 with the defaults).
countINTFrames described.