CV Generalized Hough (Shape Template)
Find any shape you can draw, even half-hidden
- image
- template
- found
- positions
- scales
- angles
- votes
- bboxes
- count
Hough circles find circles. Hough lines find lines. This finds the shape you hand it - a specific logo, a bracket, a bolt head, a logo'd badge, a stencil outline - by giving it a template image of that shape. Show it a clean crop of the thing, then search the scene.
The reason this exists rather than a template match: cv2_matchTemplate correlates raw pixel intensities, so change the lighting, the colour or the surrounding texture and your score collapses. The generalized Hough transform works from edges. It builds an R-table of the template's edge points - for each gradient direction, a list of offsets to the shape's reference point - and then every edge in the scene votes for where that reference point would have to sit. Edges survive shading, colour and background changes, and partial occlusion just means fewer votes. This is genuine classical CV, in bmad4ever's ComfyUI CV pack (bmad4ever/comfyui_cv, a fork of Gerold Meisinger's opencv-comfyui), wrapping the createGeneralizedHough* classes the auto-generated raw wrappers can't reach.
The two methods, and the one you want first
method is where the time goes.
Ballard (position only, fast) - translation. One accumulator. This is the default and it's what you should get working before anything else.
Guil (position + scale + rotation, slow) - also searches scale and rotation, and the accumulator multiplies by every scale step × every angle step. With the defaults (min_scale 0.8, max_scale 1.2, scale_step 0.1, min_angle 0 to max_angle 360 at angle_step 10) that's roughly a hundred passes over the scene. Widen it and you're into minutes. Keep the ranges tight - you usually know your part rotates within ±15° and appears at roughly one size.
The inputs that decide whether it works
template - the shape itself. Its edges become the R-table, so a clean, tightly cropped outline beats a busy crop every time. Crop to the object's silhouette; don't hand it a screenshot with a desk in it.
image - the scene to search. Converted to grayscale internally, edges computed inside with canny_low (50) and canny_high (100), which apply to both template and scene. Faint outline in either one? Lower the thresholds.
min_distance (10 px) - the minimum spacing between reported detections. Without this you get one shape reported as a cluster of near-identical hits, which is the single most common "why are there 40 detections" complaint with this family.
votes_threshold (30, Ballard only) - how many edge points must agree. It scales with the template's edge-pixel count, so a small tidy template needs a lower threshold than a big messy one. Nothing found → lower it. Everything found → raise it.
The advanced block is mostly Guil's: position_votes, scale_votes, angle_votes (100 / 100 / 10000), dp (accumulator resolution divisor - 2.0 halves it, coarser and faster and more tolerant), and levels (gradient-direction bins, 360).
Outputs
found - branch on it with an if/else node; zero detections is a valid result, not an error. positions - Nx2 centres, feed CV Draw Points. scales and angles - per-detection scale and rotation in degrees (all 1.0 and 0 for Ballard, which doesn't search them). votes - Nx3 accumulator votes for position, scale and angle; sort or threshold on column 0 to keep only the strong hits. bboxes - one {x, y, width, height, score, label} dict per detection, sized from the template scaled by scales, so it goes straight into a Draw BBoxes node. count - how many.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"
Or search ComfyUI CV in ComfyUI Manager. Python ≥ 3.12 and a recent ComfyUI on the V3 node API. workflows/32_shape_match_warp.json puts it next to matchShapes so you can compare the two approaches on the same input.
Common issues
Guil hangs. It's not hanging, it's multiplying. Shrink max_angle/min_angle first (that's the biggest term), then angle_step up, then the scale range. dp 2.0 is the next lever.
Nothing found. Raise the Canny thresholds' tolerance (lower canny_low), lower votes_threshold, and make sure the template's scale in the scene is inside the searched range for Guil. Also check the template really is an outline-shaped crop - a soft, low-contrast, busy crop gives the R-table nothing coherent.
Hundreds of hits on a textured background. Raise votes_threshold and raise min_distance. Textured surfaces produce edges in every direction, and every one of them votes.
Detections wobble by a pixel or two between runs. That's the accumulator's resolution; dp below 1.0 refines it at the cost of time.
Inputs (18)
| Name | Type | Default | Description |
|---|---|---|---|
| image | NPARRAY,IMAGE | Scene to search (converted to grayscale internally; edges are computed inside with the canny_* thresholds). Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| template | NPARRAY,IMAGE | Image of the shape to find. Its edges become the R-table, so a clean, tightly cropped outline works far better than a busy crop. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| method | COMBO | Ballard (position only, fast) | Ballard: translation only, one accumulator, fast. Guil: also scale and rotation, orders of magnitude slower - narrow min/max scale and angle before using it. |
| canny_low | INT | 501–255 | Lower Canny hysteresis threshold used for BOTH the template and the scene edges. Lower it if the shape's outline is faint. |
| canny_high | INT | 1001–255 | Upper Canny hysteresis threshold; roughly 2-3x the lower one. |
| min_distance | FLOAT | 101–1000 | Minimum distance in pixels between two reported detections - this is what stops one shape being reported as a cluster of near-identical hits. |
| votes_threshold | INT | 301–100000 | Ballard only: how many edge points must vote for a position before it counts. LOWER it if nothing is found, raise it if the scene is full of false hits. It scales with the template's edge-pixel count. |
| dpopt | FLOAT | 1.00.5–10 | Accumulator resolution divisor: 1.0 = same resolution as the image, 2.0 = half (coarser, faster, more tolerant). |
| levelsopt | INT | 36010–3600 | Number of gradient-direction bins in the R-table. |
| min_scaleopt | FLOAT | 0.800.05–10 | Guil only: smallest template scale to search. |
| max_scaleopt | FLOAT | 1.200.05–10 | Guil only: largest template scale to search. |
| scale_stepopt | FLOAT | 0.100.01–1 | Guil only: scale increment. Every extra step costs a full pass. |
| min_angleopt | FLOAT | 00–360 | Guil only: smallest rotation in DEGREES to search. |
| max_angleopt | FLOAT | 3600–360 | Guil only: largest rotation in degrees. |
| angle_stepopt | FLOAT | 10.00.1–90 | Guil only: rotation increment in degrees. 1 degree over the full circle is 360 passes - start coarse. |
| position_votesopt | INT | 1001–100000 | Guil only: vote threshold of the POSITION stage. |
| scale_votesopt | INT | 1001–100000 | Guil only: vote threshold of the SCALE stage. |
| angle_votesopt | INT | 100001–1000000 | Guil only: vote threshold of the ANGLE stage. |
Outputs (7)
| Name | Type | Description |
|---|---|---|
| found | BOOLEAN | True when at least one instance was detected - branch on it with 'Basic data handling: IfElse'. |
| positions | NPARRAY | Nx2 float32 centres of the detected shapes - feed 'CV Draw Points'. |
| scales | NPARRAY | (N,) float32 template scale per detection (all 1.0 for Ballard, which does not search scale). |
| angles | NPARRAY | (N,) float32 rotation in DEGREES per detection (all 0 for Ballard). |
| votes | NPARRAY | Nx3 int32 accumulator votes (position, scale, angle) - the detection confidence. Sort or threshold on column 0 to keep the strongest hits. |
| bboxes | BOUNDING_BOX | One {x, y, width, height, score, label} dict per detection: the template's size scaled by 'scales', centred on the position - feed the core 'Draw BBoxes'. |
| count | INT | How many instances were detected; 0 is valid, not an error. |