Nodes/comfyui-lsnet/Kaloscope Extract Features
ComfyUI Node

Kaloscope Extract Features

Turn a folder of images into vectors you can compute on

By spawner1145·Created 12 months ago·Updated 2 days ago· 104
Kaloscope Extract Features
  • image
  • model
  • features
◄output_typedefault►
◄layers-1►
◄intermediate_normtrue►

Most of this pack is built around two questions: what style is this, and what does it look like. Extract Features is the raw-material node underneath both. It runs the Kaloscope backbone over an image batch and hands you the actual numbers - a tensor - with no interpretation attached. No tags, no chart, no similarity score. Just features on the way to something else.

Reach for it when the pack's convenience nodes don't do what you want. Feed the tensor into Feature Analysis for charts, into Clustering for grouping, or into whatever custom node you've got that eats tensors. It's also the node that makes the "same features, many charts" trick possible: extract once, hang four analysis nodes off the output.

The mechanism

Straightforward and worth knowing in detail because the shapes matter. Each image in the batch is converted back to PIL, run through the checkpoint's own training transform, stacked into chunks - two images at a time, hardcoded in the pack - and pushed through the model's extraction path. Output is concatenated, cast to float32, moved to CPU, made contiguous. One features tensor comes out.

What's in that tensor depends entirely on output_type, a 19-way enum that looks intimidating and mostly isn't:

  • default - whatever the model's config.json says its feature source is (backbone pooling or projector). This is the one to use if you're just comparing images.
  • backbone / cls / mean / cls_mean - pooled global vectors, [B,D] or [B,2D]. cls_mean concatenates the class token with the mean of the patch tokens, which is the popular choice for retrieval-style work.
  • projector - the projection head's output [B,P]. Only exists if the checkpoint actually has a projector.
  • patch_tokens [B,N,D], patch_map [B,D,H,W], storage_tokens [B,R,D], all_tokens [B,1+R+N,D], prenorm - the local views. Patch maps are what the patch_energy chart needs; storage/register tokens only exist on architectures that have them.
  • Everything prefixed intermediate_ - the same menu, but per layer. These give [B,L,…], where L is how many layers you asked for, always keeping the layer axis even if you only asked for one.

layers is a string of comma-separated indices, 0-based, negatives counting from the end: -1 for the final layer, 8,9,10,11 or -4,-3,-2,-1 for the last four blocks of a ViT-B. LayerNorm is applied to intermediate features by default (intermediate_norm); the prenorm variants never apply it. Order in equals order out, and repeated indices are rejected.

The real limits

LSNet checkpoints support default and backbone only. Pick patch_tokens on an LSNet model and you get a clear "requires a DINOv3 model" error, which is at least honest - but it means the fancy output types need a DINOv3-based checkpoint, and the public Kaloscope v1/v2 releases are LSNet. The v3 DINOv3 models aren't out yet.

Selecting multiple layers only works when the selected layers have identical shapes. That's a hard error the moment you mix a ConvNeXt stage with a different resolution, so with ConvNeXt pick one stage at a time.

Two more that surprise people. The node refuses batches over 512 images, and it extracts two at a time regardless of your GPU, so a 300-image folder is a leisurely job rather than a fast one. And the output is float32 on CPU - patch_tokens on 512 images is a genuinely large tensor, big enough to be the thing that makes ComfyUI swap. If you only need global vectors, stay with default or cls_mean.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/spawner1145/comfyui-kaloscope
cd comfyui-kaloscope
python -m pip install -r requirements.txt

That pulls torch>=2.4.1, timm>=1.0.20, scikit-learn, scipy, matplotlib, safetensors and friends - plus triton-windows on Windows, which LSNet's attention kernel actually needs. Model weights go in ComfyUI/models/kaloscope/<folder>/ with config.json next to them (Kaloscope v2, or ModelScope).

Wiring

image batch → Kaloscope Extract Features → features ├→ Kaloscope Feature Analysis (chart)
model ───────┘                            (TENSOR)   └→ Kaloscope Clustering

Keep the config fixed across a comparison - same model, same output_type, same layers. Two tensors extracted with different settings are in different coordinate systems, and every downstream similarity you compute between them is nonsense that looks like a result.

CategoryKaloscope

Inputs (5)

NameTypeDefaultDescription
imageIMAGE—
modelKALOSCOPE_MODEL—
output_typeoptCOMBOdefault19 options: default, backbone, cls, mean, cls_mean, projector, +13
layersoptSTRING-1Intermediate layer indices, e.g. -1 or 8,9,10,11; negative indices count from the end.
intermediate_normoptBOOLEANtrueApply model LayerNorm to intermediate features; prenorm always skips it.

Outputs (1)

NameTypeDescription
featuresTENSOR—