Nodes/ComfyUI CV/CV VLM Generate
ComfyUI Node

CV VLM Generate

A 3B vision-language model inside cv2.dnn, one slow token at a time

By bmad4ever·Created 3 months ago·Updated 14 days ago· 1
CV VLM Generate
  • image
  • response
  • new_tokens
◄model▾►
◄promptcap en ►
◄max_new_tokens32►
◄vision_fileauto (from model folder)►
◄embed_fileauto (from model folder)►
◄lm_fileauto (from model folder)►
◄tokenizerauto (from model folder)►
◄stop_token_ids1►
◄image_size224►
◄pixel_mean0.5►
◄pixel_std0.5►
◄enginenew graph►

Show the node an image, give it a text prompt, get an answer. That's a vision-language model, and this one lives entirely inside OpenCV's DNN module - three ONNX graphs stitched together by hand instead of one transformers pipeline doing it for you.

The point is not that it's the best way to caption an image. It isn't. The point is that it's this pack's way, and the pack's whole premise is "everything goes through cv2.dnn on principle."

How it works

It mirrors OpenCV's own vlm_inference.py sample, and it expects a PaliGemma-shaped export split into three parts inside one model folder:

  • a vision encoder (SigLIP in the PaliGemma case) that turns the resized image into image-feature tokens,
  • a token embedding graph that turns prompt token ids into text embeddings,
  • a language model that takes [image features | text embeddings] and produces logits.

Then greedy decoding in a loop, in the pack's interruptible DNN worker, until a stop token or max_new_tokens. The fusion is the specific PaliGemma one: image features are concatenated before the prompt embeddings. Any VLM that fuses differently - cross-attention instead of a prefix, say - needs a different node or its own subgraph, and no amount of prompt fiddling will fix that.

Setup, which is the hard part

Models go in ComfyUI/models/llm, one folder per model. The node's model widget is a folder dropdown; it finds the three .onnx files and the tokenizer inside. The author's reference is opencv/paligemma2-3b-pt-224, a 3-part export plus an OpenCV-format config.json and a gemma2 SentencePiece tokenizer pulled from opencv_extra. The parts, the tokenizer and the preprocessing constants all have to match each other - pixel_mean/pixel_std of 0.5 are SigLIP's, 1 is Gemma's <eos> as the stop token.

Two things worth reading in the source-of-truth file before you download anything. The PaliGemma2 export is under the Gemma Terms - not an open licence, not redistributable, gated behind accepting terms on Google's model page. And the graphs need OpenCV 5's GRAPH engine; the engine widget defaults to new graph and, on OpenCV 5.1 where the two engines merged, auto resolves to the same thing.

The rest of the widgets are the usual preprocessing knobs: image_size (224 for the reference model), pixel_mean / pixel_std as one value or r, g, b, comma-separated stop_token_ids, and per-part overrides if your folder layout is unusual - vision_file, embed_file, lm_file, tokenizer, all defaulting to auto.

Outputs are simple: response, a string, and new_tokens, a count.

The honest tradeoff

CV VLM Generate runs on the CPU through cv2.dnn in fp32. The node's own tooltip says a 3B LM in full-sequence mode re-runs the whole prefix every step and takes about a minute per token, which is why max_new_tokens defaults to 32. Thirty-two tokens is a short sentence, and you're waiting half an hour for it.

Meanwhile the community already solved captioning: JoyCaption for natural-language descriptions, Florence-2 for something small and fast, WD14 for Danbooru tags. The knowledge base's llm-in-comfyui.md essay walks the whole comparison, including the multi-subject attribution weakness they all share. If your actual goal is captions for a LoRA dataset or a prompt seeded from an existing image, use one of those, on the GPU, in seconds.

So who's this for? Someone who wants to see the machinery - how a VLM's three graphs connect, what "prefix the image features" means in practice - or someone reaching for a model core doesn't support. Same caveat the pack's README applies to its DNN examples generally.

Installing it

Manager → ComfyUI CV, or:

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv

Restart. Python ≥ 3.12, a recent ComfyUI on the V3 node API, opencv-contrib-python-headless~=5.0.0.93. The pack registers models/llm with ComfyUI's folder paths itself, so the dropdown populates on a cold start. GPL-3.0, a fork of opencv-comfyui, and the author's disclaimer - LLM-written code, no support promise, verify before production - is on the record here as much as anywhere.

Where people get burned

A folder with the wrong parts in it. Missing config.json, missing tokenizer, or an export with inputs named differently from pixel_values / input_ids / inputs_embeds. auto is a search, not a magic wand; the per-part overrides exist precisely for the layout that doesn't match.

Wrong preprocessing. Feeding SigLIP an image normalized with the wrong mean makes the features garbage, and a VLM with garbage features answers confidently about nothing in particular. Check image_size, pixel_mean, pixel_std against the model card.

Not stopping. Wrong stop_token_ids means generation runs to max_new_tokens and you get a chunk of rambling that looks like a truncated answer. <eos> = 1 for the Gemma family.

Waiting. Set max_new_tokens to something small while you're wiring the graph, get the pipeline right, and only then raise it - there's nothing more demoralising than debugging a one-minute-per-token setup by running it.

Categoryimage/CV/dnn

Inputs (13)

NameTypeDefaultDescription
imageNPARRAY,IMAGEThe image to ask about (first frame of a batch). Resized to image_size x image_size. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
modelCOMBOModel FOLDER under models/llm (e.g. opencv/paligemma2-3b-pt-224). Its vision / embedding / language parts and its tokenizer are found inside it - override any of them below if the layout is unusual.
promptSTRINGcap en Task prompt ('cap en\n' = caption in English for PaliGemma; a question for instruct VLMs).
max_new_tokensINT321–1024Maximum tokens to generate. NOTE: 'full sequence' mode re-runs the whole prefix every step - a 3B fp32 LM takes ~a minute per token on CPU, keep this small.
vision_fileoptCOMBOauto (from model folder)Override the vision encoder .onnx (input 'pixel_values', e.g. PaliGemma2's SigLIP vision_model.onnx -> (1, 256, 2304) image-feature tokens). 'auto' finds it in the model folder.
embed_fileoptCOMBOauto (from model folder)Override the token embedding .onnx (input 'input_ids' -> embedding rows, e.g. embedding.onnx).
lm_fileoptCOMBOauto (from model folder)Override the language model .onnx (input 'inputs_embeds' -> logits, e.g. gemma2_3b.onnx).
tokenizeroptCOMBOauto (from model folder)Override the tokenizer folder. 'auto' uses the model folder's own config.json + tokenizer.json (PaliGemma2: the gemma2 SentencePiece pair).
stop_token_idsoptSTRING1Comma-separated ids that end generation. PaliGemma/Gemma: 1 (<eos>).
image_sizeoptINT22432–2048Square size the image is resized to for the vision encoder.
pixel_meanoptSTRING0.5Pixel normalization mean: one value or 'r, g, b' (SigLIP: 0.5).
pixel_stdoptSTRING0.5Pixel normalization std: one value or 'r, g, b' (SigLIP: 0.5).
engineoptCOMBOnew graphDNN engine for cv2.dnn.readNetFromONNX. These graphs need the OpenCV 5 GRAPH engine. OpenCV 5.1 merged its two engines into one, so 'auto' is the graph engine there and this default follows the build.

Outputs (2)

NameTypeDescription
responseSTRINGThe generated answer (stop tokens removed).
new_tokensINTNumber of tokens generated.