CV Train Classifier
Train a real classifier in the graph, then check whether it's lying to you
- features
- labels
- query_features
- predictions
- eval_labels
- accuracy
- train_accuracy
- confusion
- success
- query_predictions
- model_yaml
You have features - stacked points, HOG rows, embeddings from a model that already ran - and integer labels for them. You want a classifier. CV Train Classifier fits one from the cv2.ml family and, more importantly, tells you whether it works instead of letting you assume it does.
The design choice that matters: the model never leaves the node. Training, evaluation and the optional prediction all happen in one execute call, so ComfyUI's caching stays correct (a node that trained on the side and cached its model would give you stale predictions the moment the inputs changed, with no visual sign).
What happens in one call
Your rows get split into a train set and a held-out test set, stratified per class and deterministic per seed. The chosen algorithm fits on the train split, then scores on the test split. You get accuracy, train accuracy, and a confusion matrix. Same seed, same split, same model - which is the point. Reproducible experiments inside a graph that otherwise re-rolls seeds at you.
The inputs that matter
- features -
(N, D), one sample per row. Any shape flattens to rows, so a stacked blob of small image patches works. - labels -
(N,)integer class labels, aligned with the rows.CV Stack Feature Classesproduces exactly this shape. - algorithm - the menu is the point here. SVM is the strong default. KNearest is instance-based (no real training, just neighbours voting). RTrees/DTrees are tree ensembles, Boost needs exactly two classes, NormalBayes fits gaussians, LogisticRegression is linear, and ANN_MLP is a small fully-connected net trained from scratch by backprop - note it cannot fine-tune a pretrained network, that's not a thing
cv2.mldoes. - test_fraction (default 0.3) - the held-out share. Set it to 0 and evaluation happens on the training set, which is optimistic to the point of meaningless, though it's fine if you just want decision regions for a demo.
- query_features - optional. Wire
CV Coordinate Gridin here and the model classifies every pixel-coordinate pair, which you reshape and colour-map to paint the decision regions. That's the trick that turns a classifier into a picture.
There are per-algorithm knobs - svm_kernel, svm_tuning, svm_c, svm_gamma, knn_k, tree_depth, forest_size, mlp_hidden, mlp_iterations - plus seed. For SVM, svm_tuning = trainAuto cross-validates a grid over C and gamma: slower, and the honest default once you're past a toy.
The outputs
predictions and eval_labels are aligned arrays - compare them yourself if you want a per-sample view. accuracy is the headline number, train_accuracy is the one to compare it against: a big gap between them is overfitting in one line. confusion is a (K, K) int32 matrix, true class down the rows and predicted across the columns; upscale it with cv2_resize (INTER_NEAREST) and look at it with Preview CV Array (heatmap) rather than trusting a single accuracy figure. query_predictions is one label per query row, and model_yaml is the serialized model - real OpenCV YAML that cv2.ml.SVM_load reads back outside ComfyUI, which is how anything you train here escapes the graph.
success is false when training was impossible: no samples, a single class, or an algorithm/data mismatch like Boost with three classes. It doesn't raise - you get safe empty outputs and query predictions defaulting to the first class so your reshapes still work. Gate the downstream pretty graph on it with an if/else.
Installing it
Manager → ComfyUI CV, or:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
Restart. Python ≥ 3.12, ComfyUI on the V3 node API, opencv-contrib-python-headless~=5.0.0.93. Nothing to download. Do keep the contrib wheel: swap in plain opencv-python and you'll lose chunks of the pack silently.
Where people get burned
Gamma units. svm_gamma is scale-dependent and the tooltip is the honest answer: for raw pixel coordinates you want something like 0.001-0.01, for normalized features 0.1-10. Leave it at 1 with RBF on pixel coords and you get a boundary that's noise in your input.
Judge by confusion, not accuracy. With three imbalanced classes, 85% accuracy can mean the model predicts the big class and ignores the other two. The matrix shows you that instantly.
A held-out set that isn't held out. If your "test" features were selected using the labels you're testing against, no amount of test_fraction rescues the number.
Expecting deep features to be plug-and-play. cv2.ml classifiers are small-capacity, decades-old implementations. They'll happily fit a feature vector a modern network gave you - that's an effective and cheap setup - but the ceiling is the classifier, not the embedding.
Inputs (15)
| Name | Type | Default | Description |
|---|---|---|---|
| features | NPARRAY | (N, D) training samples, one row per sample (any shape flattens to rows): stacked points, HOG rows, deep embeddings... | |
| labels | NPARRAY | (N,) integer class label per row - e.g. the labels of 'CV Stack Feature Classes'. | |
| algorithm | COMBO | SVM (support vector machine) | cv2.ml algorithm to train. SVM is the strong default; KNearest is instance-based (no real training); RTrees/DTrees/Boost are tree-based (Boost handles exactly 2 classes); NormalBayes fits gaussians; LogisticRegression is linear; ANN_MLP is a small fully-connected net trained by real backpropagation (from scratch - cv2 cannot finetune pretrained deep nets). |
| test_fraction | FLOAT | 0.300–0.9 | Held-out fraction per class for the accuracy / confusion outputs. 0 evaluates on the training set itself (optimistic - fine for decision-region demos). |
| query_featuresopt | NPARRAY | Optional extra samples to classify with the trained model (same feature length D), e.g. an 'OpenCV Coordinate Grid' to paint decision regions. | |
| svm_kernelopt | COMBO | RBF (gaussian) | SVM only: the kernel. RBF handles curved boundaries; linear is fastest and best when D is large (e.g. HOG/deep features). |
| svm_tuningopt | COMBO | manual (use C and gamma) | SVM only: manual uses the C / gamma below; trainAuto cross-validates a grid over them (slower, robust default for real features). |
| svm_copt | FLOAT | 1.000.000001–1000000 | SVM only (manual): soft-margin penalty C. Higher fits the training set tighter (risk of overfitting). |
| svm_gammaopt | FLOAT | 1.001e-9–1000000 | SVM only (manual, RBF/poly/sigmoid/chi2): kernel width. Higher = wigglier boundary. For pixel coordinates try 0.001-0.01; for normalized features 0.1-10. |
| knn_kopt | INT | 51–256 | KNearest only: how many neighbors vote. |
| tree_depthopt | INT | 81–64 | RTrees/DTrees/Boost only: maximum tree depth. |
| forest_sizeopt | INT | 641–2048 | RTrees: number of trees; Boost: number of weak learners. |
| mlp_hiddenopt | INT | 321–4096 | ANN_MLP only: neurons in the single hidden layer. |
| mlp_iterationsopt | INT | 3001–100000 | ANN_MLP: RPROP epochs; LogisticRegression: batch gradient iterations. |
| seedopt | INT | 00–2147483647 | Seed for the stratified shuffle split (and cv2's RNG) - same seed, same split, same model. |
Outputs (8)
| Name | Type | Description |
|---|---|---|
| predictions | NPARRAY | (M,) int32 predicted label of every EVAL sample (the held-out split, or the train set when test_fraction = 0). Aligned with eval_labels. |
| eval_labels | NPARRAY | (M,) int32 ground-truth label of every eval sample - compare against predictions. |
| accuracy | FLOAT | Fraction of eval samples classified correctly (chance = 1 / class count). |
| train_accuracy | FLOAT | Accuracy on the training split itself - much higher than 'accuracy' means overfitting. |
| confusion | NPARRAY | (K, K) int32 confusion matrix over the eval set: row = true class, column = predicted. Upscale with cv2.resize (INTER_NEAREST) + 'Preview CV Array' (heatmap) to see it. |
| success | BOOLEAN | False when training was impossible (no samples, one class, algorithm/data mismatch) - gate downstream consumers with if/else. |
| query_predictions | NPARRAY | (Q,) int32 predicted label per query_features row (empty when unconnected; all first-class when success=false, so reshapes stay valid). |
| model_yaml | STRING | The trained model serialized as OpenCV YAML (cv2.ml.SVM_load & co. read it back outside ComfyUI). Empty when unavailable. |