Nodes/comfyui-lsnet/Kaloscope Feature Analysis
ComfyUI Node

Kaloscope Feature Analysis

Eighteen ways to look at your image set

By spawner1145·Created 12 months ago·Updated 2 days ago· 104
Kaloscope Feature Analysis
  • features
  • images
  • visualization
  • analysis_json
  • distance_matrix
◄chart_type▾►
◄metriccosine►
◄normalizetrue►
◄cluster_methodkmeans►
◄n_clusters3►
◄top_k2►
◄reference_index0►
◄tensor_layoutauto►
◄layer_index-1►
◄layer_poolingselected►
◄token_poolingmean►
◄labels►
◄seed42►
◄perplexity5.00►
◄dbscan_eps0.35►
◄dbscan_min_samples2►
◄max_dimensions32►
◄heatmap_ordercluster►
◄grid_width0►
◄width1400►
◄height1000►

This is the node where Kaloscope stops being a classifier and becomes a visualisation tool. You hand it a feature tensor and it draws you a picture: a network of which images are nearest neighbours, a distance heatmap, a t-SNE scatter with colour-coded clusters, a dendrogram, a silhouette plot. Eighteen chart types, all from the same tensor.

The best part is what it doesn't do: it never touches the GPU or the model. Everything here is CPU numpy, scipy and matplotlib. That means you can extract features once, then hang half a dozen Feature Analysis nodes off the same tensor and get half a dozen different views for free. It's also why the node is the one to use when you're iterating - change the metric, rerun, no inference.

What's actually in the box

The chart list splits into three rough families. Geometry: relationship_graph (MDS-laid network with edges labelled by real feature distance), pca_scatter, mds_scatter, tsne_scatter, dendrogram. Pairwise matrices: distance_heatmap, similarity_heatmap, feature_heatmap, dimension_correlation, cluster_centroid_heatmap. Stats: nearest_neighbors, distance_distribution, silhouette, cluster_sizes, pca_variance, feature_statistics, outlier_scores, patch_energy.

For a first look, relationship_graph with cluster_method=kmeans and a done-by-you n_clusters is the one that answers "what have I got here?" Drop to similarity_heatmap when you want the raw numbers visible, and outlier_scores when you suspect your reference folder is contaminated - that chart ranks images by mean distance to their neighbours, so the odd one out floats to the top.

Outputs

Three. visualization is a normal IMAGE - wire it to Preview Image or Save Image. analysis_json carries the underlying numbers: distances, similarities, neighbour rankings, cluster labels, projection coordinates, statistics. distance_matrix is the [B,B] tensor, for when something downstream wants the matrix rather than a picture of it.

Inputs you'll actually touch

features (from Extract Features) and chart_type are required. Of the long optional list, these are the ones that change your answer:

metric (cosine, euclidean, manhattan) and normalize (default on). Normalising each vector to unit length before distance and clustering is the right default for style features - it stops a high-magnitude image from looking distant to everything. Turn it off only if you care about magnitude; note that feature-statistics charts and patch energy always use the un-normalised input, which is deliberate.

cluster_method (kmeans, agglomerative, dbscan, none) plus n_clusters, dbscan_eps, dbscan_min_samples. KMeans wants Euclidean geometry and doesn't care about metric; agglomerative and DBSCAN use the metric you picked. If you don't know how many groups your images contain, DBSCAN is the honest choice - it'll also label loners as noise.

tensor_layout, and this is the field that causes most of the errors. auto reads [B,D] as vectors and [B,N,D] as tokens, but a 4D tensor is ambiguous and refuses to guess - you must say spatial for [B,D,H,W] or layer_tokens for [B,L,N,D]. The matching pairs: vectors, tokens, spatial, layer_vectors, layer_tokens, layer_spatial, or flatten to squash everything after the batch dimension. For intermediate extractions, layer_index picks which layer to analyse (layer_pooling=mean averages them) and token_pooling resolves token tensors to one vector per image.

labels - one name per line or a JSON array - is worth the ten seconds it takes. Unlabelled charts are unreadable once you've got twenty images on them. reference_index picks the subject of nearest_neighbors, top_k sets the neighbour count for edges and isolation scores, images is thumbnails only (it never runs inference, so order must match your features), and width/height (512–4096) control render size. seed and perplexity matter only for t-SNE; grid_width only for patch energy, where 0 auto-infers a square grid. max_dimensions and heatmap_order tune the feature heatmap, cluster ordering groups similar dimensions together.

Install and wiring

cd ComfyUI/custom_nodes
git clone https://github.com/spawner1145/comfyui-kaloscope
cd comfyui-kaloscope
python -m pip install -r requirements.txt
image batch → Kaloscope Extract Features → features → Kaloscope Feature Analysis → Preview Image
model ───────┘                                              (TENSOR)                 ↑

You need the model and its config.json in ComfyUI/models/kaloscope/<folder>/ for extraction; the analysis node itself needs no weights, so once a features tensor exists this half of the graph is portable.

Errors you'll hit

Over 512 images per graph, and the node refuses rather than chewing through your RAM. Any NaN or infinity in the tensor, and it refuses too - usually a sign a comparison upstream divided by a zero vector. An empty dimension in the tensor gets the same treatment.

Then the ordinary statistical stuff, which isn't this node's fault but shows up here: t-SNE with fewer points than your perplexity setting, KMeans asked for more clusters than you have images, heatmaps that look like uniform mush because thirty images of one artist genuinely are one blob. That last one isn't a failure - it's the finding.

CategoryKaloscope/Analysis

Inputs (23)

NameTypeDefaultDescription
featuresTENSOR—
chart_typeCOMBO18 options: relationship_graph, distance_heatmap, similarity_heatmap, pca_scatter, mds_scatter, tsne_scatter, +12
metricoptCOMBOcosine3 options: cosine, euclidean, manhattan
normalizeoptBOOLEANtrueNormalize each image vector to unit L2 norm before distance/clustering.
cluster_methodoptCOMBOkmeans4 options: kmeans, agglomerative, dbscan, none
n_clustersoptINT31–512—
top_koptINT21–511Neighbor count for relation edges, neighbor ranking and isolation scores.
reference_indexoptINT00–511—
tensor_layoutoptCOMBOautoFor 4D tensors select spatial [B,D,H,W] or layer_tokens [B,L,N,D] explicitly.
layer_indexoptINT-1-128–127—
layer_poolingoptCOMBOselected2 options: selected, mean
token_poolingoptCOMBOmean2 options: mean, flatten
labelsoptSTRINGOne image name per line or a JSON array, matching feature batch order.
imagesoptIMAGEOptional thumbnails only; never used for model inference.
seedoptINT420–2147483647—
perplexityoptFLOAT5.000.5–100—
dbscan_epsoptFLOAT0.350.001–100—
dbscan_min_samplesoptINT21–512—
max_dimensionsoptINT321–128—
heatmap_orderoptCOMBOcluster2 options: cluster, input
grid_widthoptINT00–4096Patch grid columns; 0 infers a square grid. Spatial maps preserve their H,W.
widthoptINT1400512–4096—
heightoptINT1000512–4096—

Outputs (3)

NameTypeDescription
visualizationIMAGE—
analysis_jsonSTRING—
distance_matrixTENSOR—