Nodes/comfyui-ccip/CCIPExtractFeature
ComfyUI Node

CCIPExtractFeature

Turn an anime crop into a number fingerprint — the node that powers everything else here

By spawner1145·Created 9 months ago·Updated 9 months ago· 0
CCIPExtractFeature
  • image
  • model
  • TENSOR
◄size384►

CCIPExtractFeature is the engine room of this pack. Feed it one anime character image and it returns a feature vector - a list of numbers that fingerprints that character, not that exact picture. Same character in a different pose, lighting, or art style gets a similar fingerprint; a different character gets a different one. That's the whole trick CCIP (Contrastive Anime Character Image Pre-Training) is built around.

If you've used IP-Adapter FaceID or PuLID for generating a consistent character, this is the mirror-image job: those inject identity into the diffusion process, CCIP just measures it. The model comes from deepghs's imgutils project and is trained on a big dataset of anime character images, so it's tuned for stylized 2D art rather than photos.

How it actually computes

Internally it's two ONNX models working together, and this node is the first half. The image gets resized to size (384 by default - that's what the models were trained on, so leave it unless you know why you're changing it), normalized with the same fixed mean/std the model expects, and pushed through model_feat.onnx. Out comes the embedding, wrapped as a TENSOR.

Then the difference and same nodes reuse that pipeline: CCIPDifference extracts a feature from each of two images and runs them through model_metrics.onnx to get a distance; CCIPSame compares that distance against a threshold. So think of this node as exposing the extract step on its own - useful when you want the embedding itself rather than a comparison.

What you'd actually use it for

The honest take: within ComfyUI, the embedding alone doesn't wire into much. The pack's other nodes take images directly, not TENSORs - so the main reasons to reach for CCIPExtractFeature are:

  • Saving features to disk (.npy) so you can run your own clustering later, in Python, without recomputing them each run
  • Building a reference library: extract once per character, then compare new images against the library offline
  • Experimenting - seeing what size and which model folder change the embedding

The input side is three fields: image (a standard IMAGE, or any image tensor - the node handles both), model (the CCIP_MODEL from CCIPModelLoader), and size. Output is a single TENSOR.

The gotcha that matters

CCIP was trained on single-character images. One character, centered, no crowd. Give it a full scene with three characters and the embedding gets mushy - it's not doing detection or segmentation for you. If you're working from video, crop the character first (person/head detection nodes or plain manual crops) and only then extract. Feed it clean crops and the difference scores stay in the honest 0.1–0.4 range; feed it scenes and everything looks vaguely similar to everything, which is useless.

Install is the pack-wide story: clone spawner1145/comfyui-ccip into custom_nodes/, restart, and make sure onnxruntime got installed from the pack's requirements. No models folder, no features - CCIPExtractFeature will throw an import or file error the moment it runs.

CategoryCCIP

Inputs (3)

NameTypeDefaultDescription
imageIMAGE—
modelCCIP_MODEL—
sizeoptINT384—

Outputs (1)

NameTypeDescription
TENSORTENSOR—