Nodes/comfyui-lsnet/Kaloscope Common Features
ComfyUI Node

Kaloscope Common Features

One vector that stands for a whole artist folder

By spawner1145·Created 12 months ago·Updated 2 days ago· 104
Kaloscope Common Features
  • reference_images
  • model
  • common_features

A style isn't one image. That's the problem this node solves. If you have fifteen paintings by the artist you're trying to imitate, comparing a new render against each of them individually gives you fifteen numbers and no verdict. Kaloscope Common Features collapses the whole folder into a single vector - the average of every image's features - so you can ask "is my render close to this artist" once instead of fifteen times.

That's the entire node. It's the boring plumbing piece of the pack, and it's the one that makes group comparisons possible.

How it works

In: a batch of reference_images and a loaded model. Out: one common_features tensor of shape [D], where D is the model's feature dimension. Under the hood it runs the backbone over each reference image (first frame only if you hand it a batch per slot), collects the pooled global feature vectors via the model's return_features path, and takes a plain arithmetic mean across them. If no references show up at all it returns a zero vector rather than erroring, which is worth knowing when a downstream similarity comes back as NaN - a zero vector has no direction to take a cosine against.

Note the consequence of "mean": everything shared across the set survives, everything that varies gets averaged toward the middle. That's good when the images share a style and differ in subject. It's bad when they don't - throw in a couple of outliers (a rough sketch, a photo, a totally different franchise) and the mean drifts toward a muddy centroid that resembles none of them. Curate the group first, or you're just averaging noise.

Where it fits

The doc's intended flow is a small pipeline: three groups of reference images, each through its own Common Features node, then all three mean vectors into Feature Comparison's group_1 / group_2 / group_3 alongside a query image. The output is a similarity score per group and a "closest group" verdict - artist A versus artist B versus artist C, decided for you.

The one thing you can't do with this output is clustering. Clustering needs to know about individual images, because the whole point is sorting per-image vectors into groups; feed it a set of means and you're clustering your own labels, which tells you nothing. The pack's own guidance is blunt about it: for clustering, feed per-image features from Extract Features, not group averages. Common Features is for comparing, not for sorting.

You also don't need the loaded model to stay loaded forever - the mean is a plain CPU tensor once it's out, so downstream nodes that take [D] tensors (Feature Comparison, and Feature Analysis with the right tensor layout) don't care whether the model is still in VRAM. The same feature tensor can feed several comparison nodes at once if you want to check one render against several artist groups in parallel branches.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/spawner1145/comfyui-kaloscope
cd comfyui-kaloscope
python -m pip install -r requirements.txt

Populate ComfyUI/models/kaloscope/<folder>/ with a checkpoint plus its config.json - Kaloscope v2 is the public one, ModelScope mirror if HF is blocked; v3 is unreleased.

Gotchas

Images in a single reference_images batch must be the same size - that's ComfyUI's rule for batching, not this pack's. Resize or pad first, and use lower-resolution copies if you're feeding a large set; these backbones resize anyway, so a 4K reference buys you nothing but sampler time.

Keep the feature configuration identical across the whole comparison. The mean is only meaningful in one embedding space - same checkpoint, same architecture, same pooling. A group averaged from a v2 LSNet run compared against a query scored by a DINOv3 v3 checkpoint is arithmetic between two unrelated coordinate systems, and it will happily return a confident-looking number that means nothing at all.

And keep the group sizes in the same ballpark when you're comparing several groups against each other. A three-image mean and a forty-image mean aren't equally stable, so the forty-image group has a slight built-in advantage: it sits closer to the middle of the space, and cosine similarity rewards that.

CategoryKaloscope

Inputs (2)

NameTypeDefaultDescription
reference_imagesIMAGE—
modelKALOSCOPE_MODEL—

Outputs (1)

NameTypeDescription
common_featuresTENSOR—