Nodes/ComfyUI CV/CV Image Hash Compare
ComfyUI Node

CV Image Hash Compare

Is this the same picture, or just a similar one?

By bmad4ever·Created 3 months ago·Updated 14 days ago· 1
CV Image Hash Compare
  • image_a
  • image_b
  • distance
  • similar
  • hash_a
  • hash_b
◄algorithmpHash (DCT, best all-round)►
◄threshold10.0►

A pixel diff answers "are these bit-identical?" - which is almost never the question you have. Re-save the same picture as WebP and the diff explodes; regenerate at a slightly different seed and the diff says "different" for two images that read as the same scene. Perceptual hashing is the tool for the actual question: is this the same picture, tolerant of rescaling, compression and mild edits. CV Image Hash Compare wraps cv2.img_hash and, crucially, wraps the comparison too - the free img_hash.* functions hand you a raw byte array, and a lone hash isn't a decision.

The inputs that matter

image_a and image_b are the two pictures. Note the tooltip on image_b: it may be a different size - that's the entire point, and the reason this doesn't just fail when someone hands it a downscaled copy. Both take frame 0 of a batch when you wire an IMAGE straight in.

algorithm is where the choice lives, and the descriptions come from the node itself:

  • pHash (DCT, best all-round) - the default. Robust across the usual edits.
  • average (fastest) - cheapest, and brightness-sensitive. Fine for near-identical frames, poor for anything graded differently.
  • block mean - the block-mean variant of the same idea.
  • colour moment (rotation-tolerant) - survives rotation.
  • Marr-Hildreth (edge structure) - the most sensitive to structural edits, so it's the one to use when you want to notice a real change.
  • radial variance (rotation-tolerant) - also survives rotation.

threshold is the distance at or below which the images count as similar, and it defaults to 10. Tune it per algorithm. The bit-based hashes report a Hamming distance in bits - roughly 2 for a rescaled copy, ~10 for the same scene, >20 for unrelated pictures, which is where that default came from. The colour-moment and radial-variance hashes report small floating-point distances instead, so a threshold near 1 is the sane range there, not 10. That scale mismatch is the single most common way people mis-configure this node; the tooltips say so explicitly.

One detail worth appreciating: the distance output always reads the same way - 0 = identical, larger = more different - for every algorithm. cv2 itself returns a similarity for radial variance (1.0 for identical), so the node inverts it for you. Comparisons don't silently flip meaning when you swap algorithms.

Outputs and wiring

Four outputs: distance (FLOAT), similar (BOOLEAN, true when distance ≤ threshold), hash_a and hash_b - the raw hashes, worth keeping if you're comparing one candidate against many references without re-hashing each time.

similar is the one you branch on. Feed it into Basic data handling: IfElse and you have a real decision in the graph: skip duplicate outputs, stop a loop when a regeneration stopped changing anything, or route a result to whichever reference it resembles. That's the whole value proposition - a perceptual hash is the only way to say "close enough" mechanically instead of with your eyes.

Where people reach for it: batch triage of generations (near-duplicate detection after a seed sweep), verifying that an upscale or a re-encode didn't actually change the picture, and matching a low-res thumbnail back to its source. It's not a similarity search engine - there's no index here - but chaining the stored hash_a against a batch of candidates is a perfectly good poor man's version.

Install

ComfyUI Manager: search the pack title comfyui_cv (bmad4ever/comfyui_cv). Manual route:

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv

Restart ComfyUI. It needs Python ≥ 3.12 and a recent ComfyUI on the V3 node API - the pack declares everything through io.ComfyNode/io.Schema with no NODE_CLASS_MAPPINGS, so a stale install won't load it. One dependency, and it's the contrib build on purpose:

pip install "opencv-contrib-python-headless~=5.0.0.93"

The version is pinned because behaviour is curated against it. If you install any non-contrib OpenCV wheel later, the shared site-packages/cv2 gets overwritten and the contrib submodules quietly empty out - the contrib nodes then vanish from the menu with nothing in the log. tools/repair_opencv_contrib.py --check and --apply in the pack both diagnose and fix that.

Final note, from the README rather than from me: this is a personal project with heavy LLM involvement, sample-tuned examples, no support promises, and a stated warning that it shouldn't go to production without independent review. Hashing is deterministic and boring in the best way, but the threshold you pick is a judgement call - verify it on your own images rather than trusting the default.

Categoryimage/CV/contrib

Inputs (4)

NameTypeDefaultDescription
image_aNPARRAY,IMAGEFirst image (3-channel). Frame 0 of a batch. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
image_bNPARRAY,IMAGESecond image. It may be a different SIZE - that is the point of a perceptual hash. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
algorithmCOMBOpHash (DCT, best all-round)pHash is the robust default. 'average' is fastest but brightness-sensitive. 'colour moment' and 'radial variance' survive ROTATION. 'Marr-Hildreth' is the most sensitive to structural edits.
thresholdFLOAT10.00–1000Distance at or below which the images count as similar. For the bit-based hashes this is in bits (~5 = near-identical, ~10 = same scene, >20 = unrelated); for 'colour moment' use a much smaller value (~1).

Outputs (4)

NameTypeDescription
distanceFLOATHow far apart the two hashes are, ALWAYS in the same direction: 0 = identical, larger = more different. The useful scale depends on the algorithm (see the node description).
similarBOOLEANTrue when distance <= threshold. Feed it to a 'Basic data handling: IfElse' to pick between two branches of the graph.
hash_aNPARRAYRaw hash of image_a - keep it to compare against many candidates without re-hashing.
hash_bNPARRAYRaw hash of image_b.