Nodes/ComfyUI CV/CV Quality Compare (cv2.quality)
ComfyUI Node

CV Quality Compare (cv2.quality)

Stop eyeballing your upscaler A/B tests

By bmad4ever·Created 4 months ago·Updated 15 days ago· 1
CV Quality Compare (cv2.quality)
  • reference
  • image
  • score
  • higher_is_better
  • quality_map
◄metricSSIM (structural similarity)►

Upscaler comparisons are a genre of internet argument built almost entirely on vibes and screenshot crops. Here's the boring, useful alternative: score two images against a reference with a real metric, get a number, and let the graph pick a winner. CV Quality Compare (cv2.quality) does that for one image; its batch sibling does it for a whole set of candidates. It's the difference between "this one looks sharper" and "this one scores 0.94 SSIM, that one 0.91."

The three inputs

reference is your ground truth - the original, the pre-filter image, the source. image is the candidate being judged. They must be the same size; that's a hard requirement, not a suggestion, and it's the first thing to check when nothing works.

metric is the choice that matters, and the node's own tooltips are the clearest short guide:

  • SSIM (structural similarity) - the default, 0–1, 1 = identical. Tracks perceived structural change, which is why it's the usual pick.
  • PSNR - log-scale error in decibels. Above ~40 dB is typically indistinguishable.
  • MSE - raw squared error. Crude, but honest and hard to fool.
  • GMSD - compares gradient structure, and correlates well with perceived sharpness loss.

Two outputs make it usable without hardcoding anything: score (averaged over the colour channels) and higher_is_better - true for SSIM/PSNR, false for MSE/GMSD. Wire that boolean into a comparison node and your workflow can pick the better of two candidates without you remembering which direction the metric runs. The third output, quality_map, is the per-pixel error map; preview it with Preview CV Array (normalize or heatmap it) and you can see where they disagree, which is usually more informative than the scalar. One quirk worth knowing: GMSD's map comes back at half the input's resolution - that's the algorithm, not a bug.

The scalar alone will mislead you, by the way. A denoiser can win on PSNR by smearing detail; a sharpener can win on GMSD while amplifying noise. Use one metric for the decision and the map for the sanity check.

What it's actually for

Two things, honestly. First, tuning: run your pipeline with two filter settings and compare against the same source. Second - and this is the trap - be very careful about using it to declare a generated image "good". There's no ground truth for a diffusion output, so comparing a re-render against the original doesn't measure quality, it measures difference, and it will happily reward the worse image for being more different. Where the metric shines is deterministic processing: resampling, denoising, compression, sharpening, contrast, upscaling kernels. Compare those against the untouched original and the number means something.

If you're scoring many candidates against one reference, don't loop this node - use CV Quality Compare Batch, which computes the reference's state once and returns the winner's index. That's the version you want in a candidate-selection graph.

Install

ComfyUI Manager, search the pack title comfyui_cv (bmad4ever/comfyui_cv). Manual:

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv

Restart ComfyUI after. It wants Python ≥ 3.12 and a recent ComfyUI on the V3 node API - the pack is all io.ComfyNode/io.Schema and carries no NODE_CLASS_MAPPINGS, so an older build won't load it. Its single dependency:

pip install "opencv-contrib-python-headless~=5.0.0.93"

The pin is the curated reference version. Stay on a contrib wheel: all four OpenCV distributions write to the same site-packages/cv2, so installing a non-contrib build over a contrib one silently empties the contrib submodules and contrib nodes (this family included) vanish from the node menu with nothing logged. tools/repair_opencv_contrib.py --check and --apply exist in the pack for exactly that.

Getting burned

Size mismatch error. Crop/scale both to the same dimensions first. This is 90% of first attempts.

Interpreted the score backwards. MSE and GMSD are lower-is-better, and a raw error of "0.03" tells you nothing until you know which metric produced it. Use the higher_is_better output instead of remembering.

Scored a comparison that isn't apples-to-apples. Two different upscalers at two different scales, or a candidate that was JPEG-compressed in between, and you're measuring the wrong variable.

And the pack's own caveat, which the README states first and loudest: heavy LLM use in the code, sample-specific tuning in several example pipelines, no support guarantees, and a stated recommendation not to run it in production without reviewing the source yourself. A metric node is a good place to apply that scepticism - you can verify SSIM against any other implementation in five minutes, and you should.

Categoryimage/CV/quality

Inputs (3)

NameTypeDefaultDescription
referenceNPARRAY,IMAGE,MASKThe ground-truth / original image. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
imageNPARRAY,IMAGE,MASKThe processed image to score against the reference. Must be the same size. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
metricCOMBOSSIM (structural similarity)SSIM tracks perceived structural change (0-1, 1 = identical) and is the usual choice. PSNR is a log-scale error in dB (>40 is typically indistinguishable). MSE is the raw squared error. GMSD compares gradient structure and correlates well with perceived sharpness loss.

Outputs (3)

NameTypeDescription
scoreFLOATOverall score, averaged over the colour channels.
higher_is_betterBOOLEANTrue for SSIM/PSNR, False for MSE/GMSD. Wire it into a comparison node so a workflow can pick the better of two candidates without hardcoding the metric's direction.
quality_mapNPARRAYPer-pixel quality/error map - preview it with 'Preview CV Array' (normalize/heatmap) to see WHERE the two images differ. GMSD's map is HALF the input's resolution.