Nodes/ComfyUI CV/CV Deep Features
ComfyUI Node

CV Deep Features

Freeze any ONNX backbone and turn it into a feature extractor

By bmad4ever·Created 4 months ago·Updated 15 days ago· 1
CV Deep Features
  • image
  • features
  • dims
  • count
  • success
◄model▾►
◄layer_name►
◄input_size224►
◄scale0.0039►
◄swap_rbtrue►
◄poolingglobal average pool (embedding)►
◄mean_r123.675►
◄mean_g116.280►
◄mean_b103.530►
◄std_r0.229►
◄std_g0.224►
◄std_b0.225►

CV Deep Features runs an ONNX backbone over your frames and gives you one embedding row per frame. That's it. No training, no gradients, no new model weights - a frozen network used as a ruler.

Why would you want that in ComfyUI? Two real reasons. First, image similarity without a model download: embed a folder of references, embed a generated image, sort by cosine distance, and you have "find me the ones that look like this" without a CLIP or a DINOv2 in sight. Second, training a small classifier on your own frames - the pack's CV Train Classifier takes these embeddings and fits a cv2.ml model, which is the classical-CV route to "is this a defect / a face / my cat".

The honest framing: cv2.dnn cannot backpropagate. Nothing here is a LoRA, nothing here touches your diffusion model, and for most style- or content-matching jobs a CLIP vision embed does more. Reach for this when you want small, deterministic and CPU-friendly, or when you want to see how the classical pipeline fits together.

How it works, and the one setting that matters

The frames go through blobFromImage (resize to a square, per-channel mean subtraction, scale, BGR→RGB) and are pushed through the net. Crucially, you pick which layer to tap. layer_name should be an interior bottleneck - for SqueezeNet 1.1 that's squeezenet0_concat7, which yields 512-dim embeddings. Leave it blank and you get the model's final output, i.e. class logits, which transfer considerably worse than an interior activation. Then the activation map is pooled into a single row: average over space per channel (position-invariant, the standard embedding) or flatten the whole thing if you want layout retained.

The node pins the net to OpenCV's classic DNN engine, because the newer fused graph engine mis-handles interior taps. Classic engine layer names carry an onnx_node! prefix, which the node adds for you when the bare name isn't found.

Inputs and outputs

Required: image (a whole IMAGE batch gives one row per frame; a plain NPARRAY counts as one frame), model (an .onnx from ComfyUI/models/onnx), layer_name, input_size (224 for the ImageNet backbones), scale (default 1/255), swap_rb (leave on - OpenCV frames are BGR, the backbones want RGB), and pooling.

Optional, and only worth touching when your backbone isn't ImageNet-trained: mean_r/mean_g/mean_b (default the ImageNet means, 0 disables) and std_r/std_g/std_b (default the ImageNet stds, 1 disables).

Outputs: features - the (N, D) float32 embedding matrix, which wires straight into CV Train Classifier or CV Stack Feature Classes - plus dims (D), count (frames embedded) and success.

Install

Part of comfyui_cv (bmad4ever/comfyui_cv) - search "ComfyUI CV" in ComfyUI Manager, or:

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"

Restart afterwards. Python ≥ 3.12 and a V3-API ComfyUI. Models are not bundled - download the ONNX backbone yourself and drop it in ComfyUI/models/onnx; the pack's model_sources.txt records URLs and licences for the ones its workflows use.

Where people get burned

  • success = false and an empty (0, 0) array. A wrong layer name or an empty batch does that, deliberately, instead of raising. Gate anything downstream with if/else on success. Also note this node can only return the batch it was given - grey or BGR NPARRAYs are treated as one frame.
  • It raises on a build with no classic engine. OpenCV 5.1 removed that engine, and tapping an interior layer becomes impossible - the call fails loudly rather than quietly handing you logits. The pinned 5.0.0.93 wheel is the supported configuration; that's not laziness, it's the one the pack is tested against.
  • Installing a plain opencv-python wheel. All four OpenCV distributions share one site-packages/cv2, so a non-contrib wheel overwrites a contrib one and the contrib nodes disappear. Fix with python tools/repair_opencv_contrib.py --check then --apply. This is the single most common way a working CV install breaks in the wild - the "cv2 DLL load failed / nodes won't load" threads are almost all this class of conflict.

Practical tip: keep the batch small. This runs through cv2.dnn on the CPU, so 200 frames of 224x224 is fine and 200 frames of 1024 is your evening gone.

Categoryimage/CV/ml

Inputs (13)

NameTypeDefaultDescription
imageNPARRAY,IMAGEFrames to embed: an IMAGE batch gives one embedding row per frame; an NPARRAY (gray/BGR/BGRA) is a single frame. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
modelCOMBOBackbone .onnx from ComfyUI/models/onnx (e.g. squeezenet1.1-7.onnx).
layer_nameSTRINGLayer to tap, e.g. squeezenet0_concat7 (the 'onnx_node!' classic-engine prefix is added automatically if needed). Blank = the model's final output - usually class LOGITS, which transfer worse than an interior bottleneck.
input_sizeINT22416–2048Square net input in px (224 for the ImageNet backbones); frames are resized by blobFromImage.
scaleFLOAT0.00390–1000000Pixel scale factor applied after mean subtraction; 1/255 = 0.0039215... maps 0-255 to the 0-1 range ImageNet models expect.
swap_rbBOOLEANtrueSwap R and B: OpenCV frames are BGR, most ONNX backbones want RGB - leave on.
poolingCOMBOglobal average pool (embedding)How a (C, H, W) activation map becomes one row: average each channel over space (C dims, position-invariant - the standard embedding), or flatten everything (C*H*W dims, keeps layout).
mean_roptFLOAT123.675-255–255RED channel mean subtracted before scaling. Default = ImageNet 0.485 * 255. Set 0 to disable.
mean_goptFLOAT116.280-255–255GREEN channel mean (ImageNet 0.456 * 255).
mean_boptFLOAT103.530-255–255BLUE channel mean (ImageNet 0.406 * 255).
std_roptFLOAT0.2290.000001–255RED channel std the scaled pixels are divided by (ImageNet 0.229). Set 1 to disable.
std_goptFLOAT0.2240.000001–255GREEN channel std (ImageNet 0.224). 1 disables.
std_boptFLOAT0.2250.000001–255BLUE channel std (ImageNet 0.225). 1 disables.

Outputs (4)

NameTypeDescription
featuresNPARRAY(N, D) float32 embeddings, one row per frame - feeds 'CV Train Classifier' / 'CV Stack Feature Classes'. Empty (0, 0) on failure.
dimsINTEmbedding length D (512 for SqueezeNet's squeezenet0_concat7).
countINTFrames embedded.
successBOOLEANFalse when the layer name does not exist in the model or no frame could be embedded - gate downstream training with if/else.