CV Seq2Seq Generate
Florence-2 and ViT-GPT2 driven through cv2.dnn
- image
- response
- new_tokens
An image goes in, text comes out. That's the node. It runs ONNX exports of encoder-decoder models - ViT-GPT2 for straight captioning, Florence-2 for the task-token style - through OpenCV's DNN module, in a graph, with no PyTorch VLM loader and no Ollama server.
The honest framing before the how: this is the heretic way to caption an image, and the pack's author says so about his own DNN examples. Core ComfyUI and the popular captioner nodes run these models natively in PyTorch on the GPU. This runs them through cv2.dnn, which in practice means CPU and one image per execution. If you want captions for a LoRA dataset, use JoyCaption or a Florence node and be done. If you want to understand what the pieces actually are - or you're already in a cv2 pipeline and don't want to drag in a second runtime - this node exists and it works.
The two shapes it supports
The description lays out two wiring shapes, and which one you're in decides half the inputs:
(a) Image-captioning style. The encoder takes the image directly - pixel_values in, hidden states out - and the decoder generates token ids against those hidden states. ViT-GPT2 is this shape. The prompt input is ignored here, because there's no text prompt to speak of.
(b) Florence-2 style. A separate vision encoder turns the image into features, those features get concatenated with the embedded text prompt as the text encoder's inputs_embeds, and the decoder takes embedded start tokens rather than raw ids. That's why there are vision_file and embed_file inputs at all: shape (b) needs an extra ONNX file for each.
Greedy decoding, in the pack's interruptible DNN worker - so a long generation doesn't lock the whole ComfyUI process with an uninterruptible C call.
The inputs that actually matter
- image - the picture. First frame of a batch; resized to a square
image_size. IMAGE/MASK or NPARRAY. - model - a folder under
ComfyUI/models/llm, not a file. Its encoder and decoder.onnxare found inside it. - prompt - the encoder's text: Florence-2's task token such as
<CAPTION>, BPE-tokenized as plain text. Ignored in shape (a). Append2viaprompt_suffix_idsfor the BART</s>. - start_token_ids / stop_token_ids - comma-separated ids. ViT-GPT2 uses
50256; Florence-2/BART uses2. Wrong values here are the single most common reason you get one token, or an empty string, or a paragraph of rambling. - max_new_tokens - 1 to 1024, default 32.
Then the shape-specific ones: vision_file and embed_file for Florence-2, tokenizer (auto uses the model folder's own config; ViT-GPT2 ships no tokenizer, so you point this at the gpt2 folder), image_size (224 for ViT-GPT2, 768 for Florence-2), and the normalization pair pixel_mean / pixel_std - Florence-2 wants 0.485, 0.456, 0.406 and 0.229, 0.224, 0.225, ViT-GPT2 wants 0.5 on both. Get those wrong and the caption is fluent nonsense, which is worse than an error.
engine defaults to the OpenCV 5 GRAPH engine, which these graphs need. Auto-encoder overrides - encoder_file, decoder_file - are there for unusual folder layouts; auto prefers the merged with-past decoder export, which is what KV-cache needs.
Outputs are response (the text, stop tokens stripped) and new_tokens. Wire response wherever a STRING goes - a CLIP text encode, a save-text node, an LLM rewriter in the same graph.
Install and models
ComfyUI Manager → ComfyUI CV → install → restart:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"
Python ≥ 3.12, recent ComfyUI on the V3 node API. Models are not bundled. You need ONNX exports of the encoder and decoder (plus the vision encoder and embedding table for shape (b)) in ComfyUI/models/llm/<your-folder-name>, with a tokenizer.json alongside; model_sources.txt at the repo root has the URLs and licences. Note the pack's own caveat: a valid ONNX export is not guaranteed to be loadable here - modern architectures are limited by both cv2.dnn and the pinned OpenCV 5.0.0.93.
Common issues
- Empty output. Almost always the id triple:
start_token_ids,stop_token_idsandprompt_suffix_idsmust match the model family. The tooltips carry the right values for both families; copy them. - One token and it stops. Stop token id is an id the model emits immediately, often because it's also the start token and you inverted which is which.
- Fluent garbage. Normalization.
pixel_mean/pixel_stdandimage_sizeare per-model constants, not preferences. Florence-2 at 224 with ImageNet stats is the classic pair of mistakes. - Nothing at all, or "graph engine" complaints. The default
engineis deliberately the graph engine; these graphs don't run on the old one. - It's slow, and that surprises you. It's
cv2.dnnon CPU. That's the cost of doing it the OpenCV way, and it's the reason the pack's author tells you to use the core node when one exists.
Inputs (17)
| Name | Type | Default | Description |
|---|---|---|---|
| image | NPARRAY,IMAGE | The image to describe (first frame of a batch). Resized to image_size x image_size. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| model | COMBO | Model FOLDER under models/llm (e.g. opencv/vit-gpt2-image-captioning). Its encoder and decoder .onnx are found inside it - override them below if the layout is unusual. Note ViT-GPT2 ships no tokenizer of its own: set 'tokenizer' to the gpt2 folder. | |
| prompt | STRING | Text fed to the ENCODER (task instruction, e.g. Florence-2's '<CAPTION>' - BPE-tokenized as plain text exactly like HF does, add suffix '2' for the BART </s>). Ignored in shape (a). | |
| start_token_ids | STRING | 50256 | Comma-separated DECODER start ids (ViT-GPT2: 50256 = <|endoftext|> BOS; Florence-2: 2). |
| max_new_tokens | INT | 321–1024 | Maximum tokens to generate. |
| encoder_fileopt | COMBO | auto (from model folder) | Override the text/main encoder .onnx. Shape (a): input 'pixel_values' (ViT-GPT2's encoder_model.onnx). Shape (b): inputs 'inputs_embeds' + 'attention_mask' (Florence-2's encoder_model.onnx). |
| decoder_fileopt | COMBO | auto (from model folder) | Override the decoder .onnx (-> logits; inputs 'input_ids' or 'inputs_embeds' + 'encoder_hidden_states'). 'auto' prefers the merged with-past export, which is what KV-cache needs. |
| vision_fileopt | COMBO | [none] | [none] = shape (a): the image goes straight into the encoder. Otherwise a separate vision encoder .onnx ('pixel_values' -> image features, e.g. Florence-2's vision_encoder.onnx) whose output is prefixed to the embedded prompt; 'auto' finds it in the model folder. |
| embed_fileopt | COMBO | [none] | [none] = the decoder takes raw 'input_ids'. Otherwise a token-embedding .onnx ('input_ids' -> embedding rows, e.g. Florence-2's embed_tokens.onnx) used for the encoder prompt and the decoder steps. |
| tokenizeropt | COMBO | auto (from model folder) | Override the tokenizer folder. 'auto' uses the model folder's own config.json + tokenizer.json - ViT-GPT2 ships none, so point this at the gpt2 folder. |
| prompt_token_idsopt | STRING | Comma-separated ids fed to the encoder INSTEAD of encoding 'prompt' - for special/task tokens the OpenCV tokenizer cannot produce. Blank = encode 'prompt' as plain text. | |
| prompt_suffix_idsopt | STRING | Comma-separated ids APPENDED to the encoded prompt (BART-style tokenizers end the encoder input with </s>: use '2' for Florence-2). | |
| stop_token_idsopt | STRING | 50256 | Comma-separated ids that end generation (ViT-GPT2: 50256; Florence-2/BART-style: 2). |
| image_sizeopt | INT | 22432–2048 | Square size the image is resized to (ViT-GPT2: 224; Florence-2: 768). |
| pixel_meanopt | STRING | 0.5 | Pixel normalization mean: one value or 'r, g, b' (ViT-GPT2: 0.5; Florence-2: 0.485, 0.456, 0.406). |
| pixel_stdopt | STRING | 0.5 | Pixel normalization std: one value or 'r, g, b' (ViT-GPT2: 0.5; Florence-2: 0.229, 0.224, 0.225). |
| engineopt | COMBO | new graph | DNN engine for cv2.dnn.readNetFromONNX. These graphs need the OpenCV 5 GRAPH engine. OpenCV 5.1 merged its two engines into one, so 'auto' is the graph engine there and this default follows the build. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| response | STRING | The generated text (stop tokens removed). |
| new_tokens | INT | Number of tokens generated. |