Nodes/ComfyUI CV/cv2.matchTemplate
ComfyUI Node

cv2.matchTemplate

You get a score map, not a match — and that's the point

By bmad4ever·Created 4 months ago·Updated 17 days ago· 1
cv2.matchTemplate
  • image
  • templ
  • mask
  • nparray
◄methodTM_CCOEFF_NORMED►

Template matching is the oldest trick in computer vision: slide a small picture across a big one, score every position, keep the best. cv2.matchTemplate is that, and it's still genuinely useful - for finding a repeated UI element, locating a watermark, snapping a crop to a known patch, or detecting one specific object you have an exact reference for. No model, no training, milliseconds.

The number one confusion is the output. You don't get a match. You don't get coordinates. You get a score map - a float image exactly (image_height - templ_height + 1) by (image_width - templ_width + 1) - and it's on you to find the peak in it. The pack leaves it raw on purpose: the code notes that matchTemplate's result is deliberately excluded from the format-echoing set, because echoing a float score map back as an IMAGE would min-max normalize it and silently destroy the values.

How it works

Both image and templ accept NPARRAY, IMAGE/MASK - and uniquely in this pack, LATENT. A LATENT link is processed in latent space: frame 0 becomes a float32 [H,W,C] array with values untouched. That's a novelty worth knowing about rather than a headline feature: template matching on raw latent values isn't comparing anything you can see, and getting a peak out of it is more of a sanity check than a workflow.

Linked IMAGEs are converted to uint8 BGR 0–255, which is the natural domain for TM_CCOEFF_NORMED - the default method. That's normalized cross-correlation: output runs roughly -1 to 1, 1.0 is a perfect match, and about 0.8 is the usual "probably it" threshold. The TM_SQDIFF family inverts the sense of the score (lower is better); if your matches look like anti-matches, you're on a SQDIFF method and should be reading the minimum.

The optional mask is the one most people never touch. It must be the same size as templ, single-channel or matching the template's channels, and it does what you'd hope: an 8-bit mask is treated as binary (only nonzero elements count, weighted 1), while a float mask weights elements by its actual values. That's how you match a subject against a busy template without the background dragging the score around - the tooltip is the author's, and it's accurate. Note that a LATENT mask is explicitly not allowed, even though the images are.

The rest of the chain

The node is one step of three. Peak-finding is cv2.minMaxLoc, which returns the minimum and maximum values and where each of them is. Then you branch: compare maxVal with your threshold, and take maxLoc as the top-left corner of the match - add templ's width and height for the full rectangle, which is what a BOUNDING_BOX-flavoured downstream node wants.

And do yourself a favour: preview the score map. Preview CV Array in heatmap mode with color_scale on turns the response surface into a picture, and you'll immediately see whether you have one crisp peak or thirty competing ones. That ten-second check answers "why did it match the wrong thing" before you start adjusting thresholds.

Installing the pack

ComfyUI CV (bmad4ever/comfyui_cv) - search "comfyui_cv" in ComfyUI Manager, or:

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"

Restart ComfyUI. Python ≥ 3.12, V3-API build, one pinned contrib OpenCV wheel, no model downloads. matchTemplate is core OpenCV; the preview and BOUNDING_BOX plumbing around it is the pack's own.

Where people get burned

Template bigger than the image. A hard cv2 error, and a common one, because people forget the template is a separate Load Image and easily crops it wrong. Sizes and data types have to be compatible; no scaling happens for you.

No scale or rotation invariance whatsoever. A template that's 5% larger, or rotated two degrees, scores like a different object. This is the limitation of the technique. If your subject varies in either, you need feature matching instead - the pack ships feature nodes and the docs' own demo is a LightGlue example - not a bigger threshold.

Repeated structure gives repeated peaks. Textured backgrounds, brickwork, tiled floors: the top score is often not the object you meant. Raise the threshold and check for a plateau of near-identical scores, which is the map telling you it has no idea.

A mask that isn't the template's size. Silent-ish failure into an OpenCV assert. Same size as templ, always.

Binary vs weighted results that confuse you. An 8-bit mask ignores its own values (nonzero = used); only a float mask weights by value. That distinction changes the score's meaning.

The pack-wide one: all four opencv-* wheels share one site-packages/cv2, so a non-contrib wheel installed over the contrib one empties the contrib submodules and those nodes stop registering at all. python tools/repair_opencv_contrib.py --check diagnoses, --apply repairs.

Categoryimage/CV/low-level/cv2 M

Inputs (4)

NameTypeDefaultDescription
imageNPARRAY,IMAGE,MASK,LATENTImage where the search is running. It must be 8-bit or 32-bit floating-point. A LATENT link is processed in latent space: frame 0 becomes a float32 [H,W,C] array (any channel count), values untouched. Arithmetic ops (add, multiply, etc.) also accept a full LATENT batch ({samples: [B,C,H,W]}) — the whole batch flows through when both inputs have the same batch size. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
templNPARRAY,IMAGE,MASK,LATENTSearched template. It must be not greater than the source image and have the same data type. A LATENT link is processed in latent space: frame 0 becomes a float32 [H,W,C] array (any channel count), values untouched. Arithmetic ops (add, multiply, etc.) also accept a full LATENT batch ({samples: [B,C,H,W]}) — the whole batch flows through when both inputs have the same batch size. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
methodCOMBOTM_CCOEFF_NORMEDParameter specifying the comparison method, see #TemplateMatchModes
maskoptNPARRAY,IMAGE,MASKOptional mask. It must have the same size as templ. It must either have the same number of channels as template or only one channel, which is then used for all template and image channels. If the data type is #CV_8U, the mask is interpreted as a binary mask, meaning only elements where mask is nonzero are used and are kept unchanged independent of the actual mask value (weight equals 1). For data type #CV_32F, the mask values are used as weights. The exact formulas are documented in #TemplateMatchModes. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.

Outputs (1)

NameTypeDescription
nparrayNPARRAY—