Nodes/ComfyUI CV/CV Cascade Detect
ComfyUI Node

CV Cascade Detect

Why the old Haar detector is still in the box

By bmad4ever·Created 3 months ago·Updated 14 days ago· 1
CV Cascade Detect
  • image
  • found
  • bboxes
  • boxes
  • centers
  • scores
  • count
◄cascade▾►
◄scale_factor1.10►
◄min_neighbors5►
◄min_size30►
◄max_size0►
◄equalize_histtrue►
◄min_confidence0.0►

Let's be blunt: for faces, CV YuNet Face Detect in this same pack is far more accurate, and if you want a licence-clean modern detector, MediaPipe's face tasks are Apache 2.0 and better again. So why does CV Cascade Detect exist?

Because it's the detector that needs no GPU, no ONNX runtime, and no weights you have to fetch from a repo that might disappear. A cascade is a small XML file, it runs on CPU in milliseconds, and it has been the default "find a face in a still" answer since 2001. That makes it the right tool for cheap pre-filtering - decide whether a frame is worth the expensive path - and for objects where nobody has trained a modern detector that you can grab. It's also the baseline: run it and YuNet side by side on your data once, and you'll have a much better feel for what the modern detectors are actually buying you.

How it works

A cascade of boosted Haar (or LBP) stages. Each stage is a pile of weak classifiers over simple rectangle features, and the useful part is the ordering: early stages reject obvious non-objects in a handful of operations, so real work is only spent on promising windows. It slides windows across a scale pyramid - hence scale_factor, the step between pyramid levels.

Inputs

  • image - an IMAGE batch uses frame 0; converted to grayscale internally.
  • cascade - the XML from ComfyUI/models/cascades. Each file detects exactly one object class: frontal face, eye, full body, and the LBP family as a faster alternative. Note the workflow reality here: OpenCV 5 no longer ships these files at all, so download them from the OpenCV 4.x branch (data/haarcascades, data/lbpcascades). They still load fine.
  • min_neighbors - the precision knob and the first one to touch. How many overlapping windows must agree before a detection is kept: raise it to kill false positives, lower it if real faces are missed.
  • scale_factor - 1.1 (10% growth per pyramid step) is the classic. Smaller finds more sizes and is much slower; 1.3+ is fast and skips sizes in between.
  • min_size / max_size - in pixels; 0 means no bound. Setting min_size to something sane is the single biggest speed win on large images.
  • equalize_hist - on by default, and it should be: Haar cascades were trained on histogram-equalised crops, so this is the standard pre-processing and it helps most on dim or unevenly lit photos.
  • min_confidence - drops detections below a stage weight, which is a nicer knob than min_neighbors for trimming a noisy result.

Outputs

found (branch on it), bboxes (core BOUNDING_BOX → Draw BBoxes), boxes (N×4 int32 raw), centers (N×2 float32, good for a Kalman step or a points overlay), scores (the cascade's stage weights) and count. Zero detections is a valid result, not an error - which matters because a face cascade on a landscape photo returning nothing is normal.

Those multiple output shapes are the useful part: bboxes composes with core nodes, boxes composes with this pack's array nodes (CV Box IoU Matrix, CV BBoxes To Array), and centers feeds tracking. Same detections, three coatings.

Install

Manager → ComfyUI CV, or:

cd ComfyUI/custom_nodes && git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"

Python ≥ 3.12, recent ComfyUI (V3 node API). The pack registers ComfyUI/models/cascades as a known folder and creates it if it's missing, so the dropdown appears empty until you drop an XML in it - then restart ComfyUI and it'll be listed.

Where people get burned

The empty dropdown. New installs have no cascades, because OpenCV 5 stopped shipping them. Grab the files from the OpenCV 4.x repo's data/haarcascades and drop them into ComfyUI/models/cascades - no conversion needed.

Expectations. Cascades are trained on frontal, upright, well-lit faces. Profiles, tilt, heavy occlusion, or a stylised/anime face will half-work or fail, and min_neighbors tuning turns "five false positives" into "two missed faces" rather than into correctness. If the detection is the point of the workflow rather than a pre-filter, use the neural node.

Categoryimage/CV/ml

Inputs (8)

NameTypeDefaultDescription
imageNPARRAY,IMAGEImage to search. An IMAGE batch uses its first frame; converted to grayscale internally. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
cascadeCOMBOCascade XML from ComfyUI/models/cascades (haarcascade_frontalface_default.xml, haarcascade_eye.xml, lbpcascade_* ...). Each file detects exactly ONE object class.
scale_factorFLOAT1.101.01–2How much the search window grows between pyramid levels. 1.1 = 10% per step: smaller finds more objects at more sizes and is much slower; 1.3+ is fast and misses objects between steps.
min_neighborsINT50–100How many overlapping detections a window must collect to be kept. This is the precision knob: raise it to kill false positives, lower it if real objects are missed.
min_sizeINT300–10000Ignore objects smaller than this many pixels on a side. 0 = no lower limit (much slower on big images).
max_sizeINT00–10000Ignore objects larger than this many pixels on a side. 0 = no upper limit.
equalize_histoptBOOLEANtrueRun cv2.equalizeHist first. Haar cascades were trained on equalized crops, so this is the standard pre-processing and usually helps on dim or unevenly-lit photos.
min_confidenceoptFLOAT0.00–1000Drop detections whose stage weight (the cascade's own confidence, from detectMultiScale3) is below this. 0 keeps everything - raise it to rank and trim without touching min_neighbors.

Outputs (6)

NameTypeDescription
foundBOOLEANTrue when at least one object was detected - branch on it with 'Basic data handling: IfElse'.
bboxesBOUNDING_BOXOne {x, y, width, height, score, label} dict per detection (score = the cascade's stage weight) - feed the core 'Draw BBoxes' node.
boxesNPARRAYNx4 int32 (x, y, w, h) - the same detections as a raw array, e.g. to seed 'CV Track Window'.
centersNPARRAYNx2 float32 detection centres - feed 'CV Draw Points' or an 'CV Kalman Filter Step'.
scoresNPARRAY(N,) float32 stage weights: how strongly the cascade voted for each detection. Higher = more confident.
countINTHow many objects were detected; 0 is valid, not an error.