CCIPDifference
How different are these two anime characters, exactly? The number is lower than you think
- image_a
- image_b
- model
- NUMBER
CCIPDifference answers one question: how similar are these two anime characters? Not "does this image match that prompt" - the actual visual identity of a drawn character. Feed it two crops of the same character and you get a small number; feed it two different characters and you get a bigger one. It's the core of the whole comfyui-ccip pack, and the output is a raw NUMBER you can do whatever you want with.
The catch, and it's a genuine trap: the output is a difference, not a similarity. Lower means more alike. If you've spent years reading cosine-similarity outputs where 0.9 is great, this flips the intuition - here 0.16 is "same character" territory and 0.40 is "clearly different people." The model card for CCIP (deepghs's Contrastive Anime Character Image Pre-Training) shows exactly that shape: same-character pairs land around 0.16–0.17, different-character pairs around 0.39–0.44.
How it works
You hand it image_a and image_b plus a model from CCIPModelLoader. Internally it extracts a feature embedding from each image (same resize-to-384 and normalize as CCIPExtractFeature, using model_feat.onnx), then runs the pair through a second model, model_metrics.onnx, which was trained to output a pairwise distance for character pairs. You get that distance back as a float. The optional size input defaults to 384 and you should leave it alone - it matches how the models were trained.
What you do with a number
That single NUMBER is oddly versatile because it's just a value in your graph:
- Compare a generated image against a reference character crop to see how close you landed
- Feed it into a math node to scale or invert it into a percentage score
- Wire it into a conditional or switch node to branch on "close enough" - though if you only need a yes/no verdict, CCIPSame already does the thresholding for you
There's no "correct" threshold baked in here, which is the flip side of the raw output being useful. The reference points are in each model folder's metrics.json - the default pruned caformer-24 model uses a threshold around 0.178 - but a single score is only meaningful relative to your images and your tolerance. Same character but wildly different art styles, different outfit, half the face cropped off? Those all inflate the difference, so don't expect 0.18 to be a hard law.
Same gotchas as the rest of the pack
The models only understand single-character crops. Whole scenes with several characters produce mush, because CCIP was trained on one-character images and does zero detection - crop first. And both inputs are plain IMAGE tensors; if you've precomputed features with CCIPExtractFeature, they don't plug into this node, which takes images only. Install is the shared story: clone spawner1145/comfyui-ccip into custom_nodes/, restart, and ensure onnxruntime is present. No local model folder under ComfyUI/models/ccip/, no difference score - the loader will tell you exactly that.
Inputs (4)
| Name | Type | Default | Description |
|---|---|---|---|
| image_a | IMAGE | — | |
| image_b | IMAGE | — | |
| model | CCIP_MODEL | — | |
| sizeopt | INT | 384 | — |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| NUMBER | NUMBER | — |