CV Match Template Multi-Scale
Template matching that survives the object being a different size
- image
- template
- template_mask
- match_crop
- score
- best_scale
- bbox
cv2.matchTemplate has one hard limitation that ruins it for real use: the template must be the same size as the thing you're looking for. Crop a logo out of one render, search a second render where it's 30% bigger, and the score collapses. Everyone's workaround is the same - loop over a range of scales, resize, match, keep the best - and everyone writes it slightly differently.
This node is that loop, packaged. Range of scales, best match wins, and because it's restricted to the normalized methods, scores stay comparable across scales. That last detail is the entire reason it works: a raw correlation score means different things at different template sizes.
How it works
image is what you search, template is what you're hunting. It also does something slightly unusual: both accept LATENT as well as IMAGE, and if you feed it a LATENT pair from the same model, the whole thing runs in latent space - coordinates then count latent cells rather than pixels, and match_crop comes back as a LATENT. Same-kind inputs only; IMAGE with IMAGE, LATENT with LATENT.
The scale sweep is scale_min → scale_max in scale_steps geometric steps (defaults: 0.25 to 2.0 in 20 steps). More steps means a finer sweep and proportionally more time, and the cost is linear in steps × image area. It's CPU work.
method defaults to TM_CCOEFF_NORMED, which is the right general choice. The other two normalized options are there, and all three give you a comparable score. template_mask lets you match a non-rectangular patch (white = compare) - but pay attention to the constraint: only TM_CCORR_NORMED supports it. Wire a mask with CCOEFF selected and you're asking for something the method can't do.
Inputs and outputs
Required: image, template, method, scale_min, scale_max, scale_steps. Optional: template_mask.
match_crop- the matched region, in whatever format you put in. IMAGE in, IMAGE crop out; LATENT in, LATENT out.score- the best confidence across all scales. Higher is better (these are the normalized methods).best_scale- the scale that won. Genuinely useful: it tells you how much bigger or smaller the thing was, which is information you can act on.bbox- a core BOUNDING_BOX you can chain intoDraw BBoxesto see what it found. It carries the score and scale as extra keys, and for a batch it's a list of per-image boxes.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
Manager → ComfyUI CV (publisher bmad4ever) if you'd rather click. Restart ComfyUI, then reload the browser tab so the node menu picks it up. Dependency is opencv-contrib-python-headless~=5.0.0.93 on Python ≥ 3.12 with a V3-API ComfyUI - no models, nothing to download.
Where it falls down
Rotation, and this is honest rather than fixable. A scale sweep doesn't rotate the template. If your object is also turned, a normalized method on a rotated template scores badly and the node returns a confident-looking wrong answer. Rotate the template beforehand, or reach for CV Generalized Hough (Shape Template) in the same pack (arbitrary-shape search) or a learned matcher.
The sweep is a search, not a solve. scale_steps=20 across 0.25–2.0 is roughly a 7% spacing - good enough to find things, coarse enough that best_scale is approximate. Narrow scale_min/scale_max around what you expect if you need precision; it's much cheaper too.
Scores near 1.0 that are wrong. The classic failure of normalized template matching on repetitive texture - the same wall tile matches everywhere. That's a property of the method, not a bug. The pack's CV NMS Boxes exists partly for the noise this node generates when you lower the threshold to catch marginal matches.
One pack-level note if this is your first CV node: this is a GPL-3.0 fork of geroldmeisinger/opencv-comfyui with a large, honest "written with heavy LLM assistance, not recommended for production without review" disclaimer at the top of its README. For template hunting in a hobby pipeline that's fine. If the match drives something you're selling, read the source - it's short.
Inputs (7)
| Name | Type | Default | Description |
|---|---|---|---|
| image | COMFY_MATCHTYPE_V3 | Image to search in. Accepts a LATENT too (then the template must be a LATENT from the same model and match_crop comes back as a LATENT). | |
| template | IMAGE,LATENT | Patch to search for. Same kind as 'image': IMAGE with IMAGE, LATENT with LATENT. | |
| method | COMBO | TM_CCOEFF_NORMED | Normalized comparison method. TM_CCOEFF_NORMED is the most robust general choice; all three give scores comparable across scales. |
| scale_min | FLOAT | 0.250.01–8 | Smallest template scale to try. |
| scale_max | FLOAT | 2.000.01–8 | Largest template scale to try. |
| scale_steps | INT | 202–100 | Number of scales tried between scale_min and scale_max (geometric spacing). More steps = more precise, slower. |
| template_maskopt | MASK | Optional mask for non-rectangular templates (white = compare). Resized along with the template at each scale. Only supported by TM_CCORR_NORMED. |
Outputs (4)
| Name | Type | Description |
|---|---|---|
| match_crop | COMFY_MATCHTYPE_V3 | The matched region, in the 'image' input's format (IMAGE in -> IMAGE crop, LATENT in -> LATENT crop). |
| score | FLOAT | Best confidence across all scales (higher = better). |
| best_scale | FLOAT | Template scale that produced the best match. |
| bbox | BOUNDING_BOX | Core BOUNDING_BOX data: per-frame lists of {x, y, width, height} dicts - compatible with Draw BBoxes, Crop By Bounding Boxes, Image Crop, etc. The match region; includes the score and scale as extra keys. A list of per-image boxes for batches. |