Kaloscope Clustering
Sort your image folder into groups
- group_1
- group_2
- group_3
- clustering_json
- visualization
You've got three hundred downloads and no idea what's in them. Which ones are the same artist? Which ones are the same kind of thing? Kaloscope Clustering is the node that answers that without any labels: feed it per-image features, tell it how many groups you expect, and it returns a label per image plus a scatter plot with the groups in colour.
It's the pack's original grouping node, and it's simpler than the Feature Analysis family - three algorithms, a handful of knobs, two outputs. Simple in a way that occasionally bites, which I'll get to.
The mechanism
Any connected group_* inputs (up to three TENSOR batches) get stacked into one [B,D] array - that's why the slots exist, so you can merge features extracted in separate passes and still cluster them together. Then:
method=kmeans runs KMeans with n_clusters and a fixed seed, and returns both labels and centers. method=dbscan runs DBSCAN on eps and min_samples, returns labels only, and can mark points as noise (label -1). method=hierarchical runs agglomerative clustering on n_clusters. Note that each algorithm ignores the other's knobs - eps and min_samples do nothing under KMeans, and n_clusters does nothing under DBSCAN.
Nothing is normalised before clustering. That's a real difference from Feature Analysis, which L2-normalises each vector by default; identical inputs can cluster differently between the two nodes. If your features are all roughly the same norm it won't matter, but if a few images dominate the scale, KMeans will happily report a cluster that is really just "the loud ones".
Inputs and outputs
Beyond method, the four you'll touch: n_clusters (2–100, default 10), eps (0.1–10, default 0.5), min_samples (1–50, default 5), and visualize with viz_method (tsne or pca) and perplexity for the t-SNE case.
clustering_json returns method, n_samples, group_sizes, and the flat labels array - plus centers for KMeans. The labels are one entry per image in the order you fed them in, and group_sizes is what tells you where each input group's slice starts and ends. visualization is an IMAGE scatter plot: t-SNE or PCA projection, points coloured by cluster, ready for Preview Image.
There's an ugly edge case worth knowing: if you connect no groups at all, or if there's only a single sample, the node doesn't error - it returns JSON with an error field and a black 64×64 image. A black preview means "you wired nothing useful," not "your clusters are boring."
The trap
n_clusters defaults to 10. KMeans and agglomerative clustering both need at least as many samples as clusters, so on a six-image test set the default blows up in sklearn with an error that points at the clustering call rather than at the default you left alone. Set it before you run. The same applies to feeding this node group averages instead of per-image features - a [D] mean vector gets stacked as a single row, so you end up clustering "one image per group" and either get an error or a confident chart of nothing. Clustering needs the per-image output of Extract Features; means belong in Feature Comparison.
For a new workflow I'd reach for Feature Analysis instead. It covers the same ground - tsne_scatter and pca_scatter with cluster_method=kmeans are this node's function - and adds cosine/Euclidean/Manhattan metrics, normalisation, per-image labels, a distance matrix output, and fifteen more charts. Clustering's advantage is that it's short: one node, one scatter, done. Use it when you want a quick look and don't care about the numbers behind it.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/spawner1145/comfyui-kaloscope
cd comfyui-kaloscope
python -m pip install -r requirements.txt
Model files go in ComfyUI/models/kaloscope/<subfolder>/ with a config.json naming the architecture - Kaloscope v2 is the public release (ModelScope mirror); v3 isn't out. The clustering itself is CPU scikit-learn, so it works fine on a machine where extraction happened elsewhere - though the features have to come from the same checkpoint to mean anything.
Last thing: the t-SNE projection is seeded, so the same features and settings give you the same picture twice. Don't read the axes as meaningful - only the neighbourhood structure is.
Inputs (10)
| Name | Type | Default | Description |
|---|---|---|---|
| method | COMBO | kmeans | 3 options: kmeans, dbscan, hierarchical |
| n_clusters | INT | 102–100 | — |
| eps | FLOAT | 0.500.1–10 | — |
| min_samples | INT | 51–50 | — |
| visualize | BOOLEAN | true | — |
| viz_method | COMBO | tsne | 2 options: tsne, pca |
| perplexity | INT | 305–100 | — |
| group_1opt | TENSOR | — | |
| group_2opt | TENSOR | — | |
| group_3opt | TENSOR | — |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| clustering_json | STRING | — |
| visualization | IMAGE | — |