Nodes/ComfyUI CV/CV Quality Compare Batch (cv2.quality)
ComfyUI Node

CV Quality Compare Batch (cv2.quality)

Score a parameter sweep against one reference

By bmad4ever·Created 4 months ago·Updated 15 days ago· 1
CV Quality Compare Batch (cv2.quality)
  • reference
  • images
  • best_score
  • best_index
  • higher_is_better
  • scores
  • quality_maps
◄metricSSIM (structural similarity)►
◄max_pixel_value255►

This is the node that turns "which of these eight settings looks best" into a number you can act on. Render N candidates into one image batch, score them all against the reference, take the winner.

Why the batch version exists

The pack also has a single-pair comparison node, which calls Quality*_compute(ref, cmp) once. That's fine for one comparison and wasteful for eight. The class API behind it - Quality*_create(ref) - precomputes the reference's state: SSIM caches the reference's mean and variance maps, GMSD its gradient magnitudes. Compare eight candidates with the one-shot call and you derive the reference eight times; use this node and it's derived once, then applied across the whole batch. Since a parameter sweep is exactly the thing that produces a batch, that's the version you want.

Picking a metric

metric has four options and they don't measure the same thing.

SSIM (structural similarity) is the default and usually the right one for perceptual comparison: it tracks structural change, scaled 0–1 with 1 meaning identical. It's forgiving of uniform brightness/contrast shifts and harsh about structural damage, which matches what your eye does.

PSNR is a log-scale error in decibels; over about 40 dB is typically indistinguishable. It's a useful, well-understood number and it is blunt - a smoothed image can score better than a subtly different one.

MSE is the raw squared error. Fine for regression checks, and per the pack's own notes the only metric here with an implementation quirk: on this OpenCV build the instance compute() binding rejects numpy images, so MSE alone falls back to the per-frame static call. Correct results, just without the cached-reference optimisation.

GMSD compares gradient structure and correlates well with perceived sharpness loss - the one to use when the question is "did this step make it softer". Note its quality map comes back at half the input resolution, because the metric works on a downsampled gradient map; don't assume a map is the same size as your image.

max_pixel_value (255) is PSNR-only and is ignored by everything else.

Inputs and outputs

reference is the ground truth - frame 0 if you feed a batch - and images is the candidate batch. Every frame must match the reference's size and channel count; that's a hard requirement, not a warning, so if you're comparing crops you need to normalise them upstream.

The outputs are designed for a workflow rather than a report. best_score and best_index give you the winner - argmax for SSIM/PSNR, argmin for MSE/GMSD, so the node has already handled the direction for you. higher_is_better is a boolean you can wire into your own comparison logic so a graph doesn't hardcode the metric's direction, which is the kind of thing that breaks quietly the day you switch metrics. scores is the (N,) array in batch order - chart it with CV Chart Series and you can see whether quality is a smooth curve over your sweep or a cliff. quality_maps is the (N,H,W,C) per-pixel error map, which is the interesting one: preview it and you see where the candidates differ, not just how much. A uniform faint map means a global shift; a bright patch means a local artefact, which is usually the information that decides your choice.

Installing

ComfyUI Manager → search "ComfyUI CV", or:

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
# restart ComfyUI

Python ≥ 3.12, a recent ComfyUI on the V3 node API, and the contrib OpenCV build: opencv-contrib-python-headless~=5.0.0.93, with numpy and torch. cv2.quality is contrib-only, and because all OpenCV wheels share one site-packages/cv2, installing a plain opencv-python over the top silently strips it - python tools/repair_opencv_contrib.py --check will tell you.

Gotchas

Metrics disagree with eyes on generative content. SSIM and PSNR reward similarity to the reference, so a denoise or upscale pass that invents pleasing detail is penalised for not matching. That's the known limitation of every full-reference metric here, and it's why "which upscale is better" is not a question SSIM answers - read the metrics as consistency scores, not quality scores, when the operation is meant to change the image.

All frames must match the reference in size. A batch with mixed resolutions will fail rather than skip the odd ones.

One sweep, one batch. Because the reference state is computed once per execution, changing the reference means changing everything downstream of it - reseed the candidates too, or you're comparing different content.

Categoryimage/CV/quality

Inputs (4)

NameTypeDefaultDescription
referenceNPARRAY,IMAGE,MASKThe ground-truth / original image (frame 0 if a batch). Its quality state is computed once. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
imagesNPARRAY,IMAGE,MASKBatch of processed candidates to score. Every frame must have the reference's size and channel count. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
metricCOMBOSSIM (structural similarity)SSIM tracks perceived structural change (0-1, 1 = identical). PSNR is a log-scale error in dB (>40 is typically indistinguishable). MSE is the raw squared error. GMSD compares gradient structure and correlates well with perceived sharpness loss.
max_pixel_valueoptFLOAT2551–65535PSNR only: the maximum per-channel pixel value used as the numerator (255 for 8-bit images). Ignored by the other metrics.

Outputs (5)

NameTypeDescription
best_scoreFLOATScore of the best frame, averaged over the colour channels.
best_indexINTIndex of the best frame in the batch - argmax for SSIM/PSNR, argmin for MSE/GMSD. Feed it to a list picker to select the winning candidate.
higher_is_betterBOOLEANTrue for SSIM/PSNR, False for MSE/GMSD. Wire it into a comparison so a workflow can rank without hardcoding the metric's direction.
scoresNPARRAYfloat32 array of shape (N,) - one score per frame, in batch order. Chart it with 'CV Chart Series'.
quality_mapsNPARRAYPer-pixel quality/error maps, float32 (N,H,W,C) - preview with 'Preview CV Array' (normalize/heatmap) to see WHERE frames differ. GMSD's map is HALF the input's resolution (the metric works on a downsampled gradient map).