CV Cascade Detect
Why the old Haar detector is still in the box
- image
- found
- bboxes
- boxes
- centers
- scores
- count
Let's be blunt: for faces, CV YuNet Face Detect in this same pack is far more accurate, and if you want a licence-clean modern detector, MediaPipe's face tasks are Apache 2.0 and better again. So why does CV Cascade Detect exist?
Because it's the detector that needs no GPU, no ONNX runtime, and no weights you have to fetch from a repo that might disappear. A cascade is a small XML file, it runs on CPU in milliseconds, and it has been the default "find a face in a still" answer since 2001. That makes it the right tool for cheap pre-filtering - decide whether a frame is worth the expensive path - and for objects where nobody has trained a modern detector that you can grab. It's also the baseline: run it and YuNet side by side on your data once, and you'll have a much better feel for what the modern detectors are actually buying you.
How it works
A cascade of boosted Haar (or LBP) stages. Each stage is a pile of weak classifiers over simple rectangle features, and the useful part is the ordering: early stages reject obvious non-objects in a handful of operations, so real work is only spent on promising windows. It slides windows across a scale pyramid - hence scale_factor, the step between pyramid levels.
Inputs
image- an IMAGE batch uses frame 0; converted to grayscale internally.cascade- the XML fromComfyUI/models/cascades. Each file detects exactly one object class: frontal face, eye, full body, and the LBP family as a faster alternative. Note the workflow reality here: OpenCV 5 no longer ships these files at all, so download them from the OpenCV 4.x branch (data/haarcascades,data/lbpcascades). They still load fine.min_neighbors- the precision knob and the first one to touch. How many overlapping windows must agree before a detection is kept: raise it to kill false positives, lower it if real faces are missed.scale_factor- 1.1 (10% growth per pyramid step) is the classic. Smaller finds more sizes and is much slower; 1.3+ is fast and skips sizes in between.min_size/max_size- in pixels; 0 means no bound. Settingmin_sizeto something sane is the single biggest speed win on large images.equalize_hist- on by default, and it should be: Haar cascades were trained on histogram-equalised crops, so this is the standard pre-processing and it helps most on dim or unevenly lit photos.min_confidence- drops detections below a stage weight, which is a nicer knob thanmin_neighborsfor trimming a noisy result.
Outputs
found (branch on it), bboxes (core BOUNDING_BOX → Draw BBoxes), boxes (N×4 int32 raw), centers (N×2 float32, good for a Kalman step or a points overlay), scores (the cascade's stage weights) and count. Zero detections is a valid result, not an error - which matters because a face cascade on a landscape photo returning nothing is normal.
Those multiple output shapes are the useful part: bboxes composes with core nodes, boxes composes with this pack's array nodes (CV Box IoU Matrix, CV BBoxes To Array), and centers feeds tracking. Same detections, three coatings.
Install
Manager → ComfyUI CV, or:
cd ComfyUI/custom_nodes && git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"
Python ≥ 3.12, recent ComfyUI (V3 node API). The pack registers ComfyUI/models/cascades as a known folder and creates it if it's missing, so the dropdown appears empty until you drop an XML in it - then restart ComfyUI and it'll be listed.
Where people get burned
The empty dropdown. New installs have no cascades, because OpenCV 5 stopped shipping them. Grab the files from the OpenCV 4.x repo's data/haarcascades and drop them into ComfyUI/models/cascades - no conversion needed.
Expectations. Cascades are trained on frontal, upright, well-lit faces. Profiles, tilt, heavy occlusion, or a stylised/anime face will half-work or fail, and min_neighbors tuning turns "five false positives" into "two missed faces" rather than into correctness. If the detection is the point of the workflow rather than a pre-filter, use the neural node.
Inputs (8)
| Name | Type | Default | Description |
|---|---|---|---|
| image | NPARRAY,IMAGE | Image to search. An IMAGE batch uses its first frame; converted to grayscale internally. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| cascade | COMBO | Cascade XML from ComfyUI/models/cascades (haarcascade_frontalface_default.xml, haarcascade_eye.xml, lbpcascade_* ...). Each file detects exactly ONE object class. | |
| scale_factor | FLOAT | 1.101.01–2 | How much the search window grows between pyramid levels. 1.1 = 10% per step: smaller finds more objects at more sizes and is much slower; 1.3+ is fast and misses objects between steps. |
| min_neighbors | INT | 50–100 | How many overlapping detections a window must collect to be kept. This is the precision knob: raise it to kill false positives, lower it if real objects are missed. |
| min_size | INT | 300–10000 | Ignore objects smaller than this many pixels on a side. 0 = no lower limit (much slower on big images). |
| max_size | INT | 00–10000 | Ignore objects larger than this many pixels on a side. 0 = no upper limit. |
| equalize_histopt | BOOLEAN | true | Run cv2.equalizeHist first. Haar cascades were trained on equalized crops, so this is the standard pre-processing and usually helps on dim or unevenly-lit photos. |
| min_confidenceopt | FLOAT | 0.00–1000 | Drop detections whose stage weight (the cascade's own confidence, from detectMultiScale3) is below this. 0 keeps everything - raise it to rank and trim without touching min_neighbors. |
Outputs (6)
| Name | Type | Description |
|---|---|---|
| found | BOOLEAN | True when at least one object was detected - branch on it with 'Basic data handling: IfElse'. |
| bboxes | BOUNDING_BOX | One {x, y, width, height, score, label} dict per detection (score = the cascade's stage weight) - feed the core 'Draw BBoxes' node. |
| boxes | NPARRAY | Nx4 int32 (x, y, w, h) - the same detections as a raw array, e.g. to seed 'CV Track Window'. |
| centers | NPARRAY | Nx2 float32 detection centres - feed 'CV Draw Points' or an 'CV Kalman Filter Step'. |
| scores | NPARRAY | (N,) float32 stage weights: how strongly the cascade voted for each detection. Higher = more confident. |
| count | INT | How many objects were detected; 0 is valid, not an error. |