CV Quality Compare (cv2.quality)
Stop eyeballing your upscaler A/B tests
- reference
- image
- score
- higher_is_better
- quality_map
Upscaler comparisons are a genre of internet argument built almost entirely on vibes and screenshot crops. Here's the boring, useful alternative: score two images against a reference with a real metric, get a number, and let the graph pick a winner. CV Quality Compare (cv2.quality) does that for one image; its batch sibling does it for a whole set of candidates. It's the difference between "this one looks sharper" and "this one scores 0.94 SSIM, that one 0.91."
The three inputs
reference is your ground truth - the original, the pre-filter image, the source. image is the candidate being judged. They must be the same size; that's a hard requirement, not a suggestion, and it's the first thing to check when nothing works.
metric is the choice that matters, and the node's own tooltips are the clearest short guide:
- SSIM (structural similarity) - the default, 0–1, 1 = identical. Tracks perceived structural change, which is why it's the usual pick.
- PSNR - log-scale error in decibels. Above ~40 dB is typically indistinguishable.
- MSE - raw squared error. Crude, but honest and hard to fool.
- GMSD - compares gradient structure, and correlates well with perceived sharpness loss.
Two outputs make it usable without hardcoding anything: score (averaged over the colour channels) and higher_is_better - true for SSIM/PSNR, false for MSE/GMSD. Wire that boolean into a comparison node and your workflow can pick the better of two candidates without you remembering which direction the metric runs. The third output, quality_map, is the per-pixel error map; preview it with Preview CV Array (normalize or heatmap it) and you can see where they disagree, which is usually more informative than the scalar. One quirk worth knowing: GMSD's map comes back at half the input's resolution - that's the algorithm, not a bug.
The scalar alone will mislead you, by the way. A denoiser can win on PSNR by smearing detail; a sharpener can win on GMSD while amplifying noise. Use one metric for the decision and the map for the sanity check.
What it's actually for
Two things, honestly. First, tuning: run your pipeline with two filter settings and compare against the same source. Second - and this is the trap - be very careful about using it to declare a generated image "good". There's no ground truth for a diffusion output, so comparing a re-render against the original doesn't measure quality, it measures difference, and it will happily reward the worse image for being more different. Where the metric shines is deterministic processing: resampling, denoising, compression, sharpening, contrast, upscaling kernels. Compare those against the untouched original and the number means something.
If you're scoring many candidates against one reference, don't loop this node - use CV Quality Compare Batch, which computes the reference's state once and returns the winner's index. That's the version you want in a candidate-selection graph.
Install
ComfyUI Manager, search the pack title comfyui_cv (bmad4ever/comfyui_cv). Manual:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
Restart ComfyUI after. It wants Python ≥ 3.12 and a recent ComfyUI on the V3 node API - the pack is all io.ComfyNode/io.Schema and carries no NODE_CLASS_MAPPINGS, so an older build won't load it. Its single dependency:
pip install "opencv-contrib-python-headless~=5.0.0.93"
The pin is the curated reference version. Stay on a contrib wheel: all four OpenCV distributions write to the same site-packages/cv2, so installing a non-contrib build over a contrib one silently empties the contrib submodules and contrib nodes (this family included) vanish from the node menu with nothing logged. tools/repair_opencv_contrib.py --check and --apply exist in the pack for exactly that.
Getting burned
Size mismatch error. Crop/scale both to the same dimensions first. This is 90% of first attempts.
Interpreted the score backwards. MSE and GMSD are lower-is-better, and a raw error of "0.03" tells you nothing until you know which metric produced it. Use the higher_is_better output instead of remembering.
Scored a comparison that isn't apples-to-apples. Two different upscalers at two different scales, or a candidate that was JPEG-compressed in between, and you're measuring the wrong variable.
And the pack's own caveat, which the README states first and loudest: heavy LLM use in the code, sample-specific tuning in several example pipelines, no support guarantees, and a stated recommendation not to run it in production without reviewing the source yourself. A metric node is a good place to apply that scepticism - you can verify SSIM against any other implementation in five minutes, and you should.
Inputs (3)
| Name | Type | Default | Description |
|---|---|---|---|
| reference | NPARRAY,IMAGE,MASK | The ground-truth / original image. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| image | NPARRAY,IMAGE,MASK | The processed image to score against the reference. Must be the same size. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| metric | COMBO | SSIM (structural similarity) | SSIM tracks perceived structural change (0-1, 1 = identical) and is the usual choice. PSNR is a log-scale error in dB (>40 is typically indistinguishable). MSE is the raw squared error. GMSD compares gradient structure and correlates well with perceived sharpness loss. |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| score | FLOAT | Overall score, averaged over the colour channels. |
| higher_is_better | BOOLEAN | True for SSIM/PSNR, False for MSE/GMSD. Wire it into a comparison node so a workflow can pick the better of two candidates without hardcoding the metric's direction. |
| quality_map | NPARRAY | Per-pixel quality/error map - preview it with 'Preview CV Array' (normalize/heatmap) to see WHERE the two images differ. GMSD's map is HALF the input's resolution. |