CV LLM Generate
A local LLM running through cv2.dnn, on purpose
- response
- full_text
- new_tokens
The name is not a lie in the usual way - but it isn't what you're hoping either. CV LLM Generate runs a decoder-only language model through OpenCV's DNN engine, generating text in the graph with no API key and no network call. What it is not is a better Ollama. It's the cv2-only path to text generation, and knowing when that's the right path is most of the decision.
What's actually happening
OpenCV 5's DNN module grew a tokenizer and KV-cache support (cv2.dnn.Tokenizer), which is what makes LLM inference possible at all in this pack: tokenization, decoding, and the generation loop all run inside the interruptible DNN worker, so a long generation doesn't lock up your session. Decoding is greedy - deterministic, no sampling, no temperature.
The model has to be an ONNX export sitting in a folder under ComfyUI/models/llm, and the pack's mirror of the OpenCV samples tells you the shapes that work: Qwen2.5 or Gemma3 instruct exports, or GPT-2. Two export routes are named - optimum-cli export onnx --task causal-lm-with-past, or the merged decoder from a transformers.js export. KV-cache mode needs the with-past / merged export; a plain causal-lm export only works in full sequence mode.
Inputs that matter
model is a folder, not a file - something like opencv/qwen2.5-0.5b-instruct - and its .onnx and tokenizer are found inside it. template wraps your prompt in the right chat format: chatml (Qwen2.5), gemma3, or none for raw completion (GPT-2). It also sets the default stop and BOS ids, which is the quiet reason to get this right rather than pasting <|im_start|> into the prompt by hand.
prompt is multiline and defaults to the entirely reasonable "What is OpenCV?". max_new_tokens defaults to 32 and the tooltip is honest about why: CPU inference is roughly a few tokens per second for a 0.5B model. Set it to 512 and go make tea.
decode_mode is the real performance switch. kv-cache is fast - one token per step through the graph - and requires the with-past export. full sequence re-feeds the entire growing sequence every step: slow, but it works with plain exports. If you have the choice, export with-past and never think about it again.
The optionals are escape hatches for unusual layouts: lm_file (which .onnx in the folder is the language model - auto picks the merged/only decoder), tokenizer (override the tokenizer folder when the export ships none of its own), stop_token_ids (comma-separated; blank uses the template default), bos_token_id (-1 = template default; Gemma3 prepends 2, the others none), and engine. That last one matters if generation fails to load: LLM ONNX graphs need OpenCV 5's GRAPH engine, and the classic 4.x engine can't parse them. On OpenCV 5.1 the two engines merged, so auto resolves to the graph engine there.
Outputs are response (generated text only - prompt stripped, stop tokens removed), full_text (prompt + response, detokenized), and new_tokens (count before stop-id removal).
Should you actually use this?
Usually not, and the pack's own README says so first: everything here goes through cv2.dnn on principle, even where ComfyUI already does the job better. For prompt enhancement or captioning, the sensible 2026 answer is an LLM tool node (Ollama-style, or a multi-provider node for the API path) rather than an OpenCV DNN wrapper - the local-LLM nodes exist because they're uncensored, free per call and offline, and that argument is about the model choice, not about routing through cv2.
Where this node is the answer: you want everything inside one process with one dependency (OpenCV), your model is already an ONNX export, and the text job is small. Then it's tidy. You also get determinism for free, which is genuinely nice when you're debugging a pipeline and don't want the wording changing under you.
Install
ComfyUI Manager, search the pack title comfyui_cv (bmad4ever/comfyui_cv). Or:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
Restart ComfyUI. Needs Python ≥ 3.12 and a recent ComfyUI on the V3 node API - the pack is written as io.ComfyNode/io.Schema with no NODE_CLASS_MAPPINGS, so an older build loads nothing. Dependency:
pip install "opencv-contrib-python-headless~=5.0.0.93"
Pinned, contrib, and yes it matters here more than elsewhere: a non-contrib wheel installed over a contrib one shares the same site-packages/cv2 and silently empties the contrib submodules, which makes contrib nodes vanish from the menu with nothing logged. tools/repair_opencv_contrib.py --check / --apply is in the pack for that.
Models are not bundled. The LLM/ONNX files are external downloads and live outside the repo; model_sources.txt at the repo root records URLs, target folders and licences. Read it before redistributing anything.
Where it breaks
A red node with an empty model dropdown. The folder isn't in ComfyUI/models/llm, or ComfyUI hasn't refetched node definitions since you put it there - reload the page.
"The classic engine cannot parse this graph." An LLM ONNX needs the graph engine; pick it explicitly, or be on OpenCV 5.1 where the engines merged.
Nothing generated, or gibberish. Almost always the template or the export type: chatml text fed to a raw-completion model (or the reverse) produces exactly the confident nonsense you'd expect, and a plain export in kv-cache mode isn't going to work at all.
It takes forever. It's CPU, it's a few tokens a second, and max_new_tokens is the dial. This is why the README points you at ComfyUI's own nodes for anything serious.
Finally, the pack's own disclaimers, which apply with full force to a DNN pipeline: heavy LLM assistance in development, test-driven overfitting risk, limited DNN support that's pinned to one OpenCV version, models you have to source and convert yourself, and a stated recommendation not to use it in production without independent review. This one is a demo you can use, not infrastructure you should depend on.
Inputs (10)
| Name | Type | Default | Description |
|---|---|---|---|
| model | COMBO | Model FOLDER under models/llm (e.g. opencv/qwen2.5-0.5b-instruct). Its .onnx and its tokenizer are found inside it; override either below if the layout is unusual. KV-cache mode needs a with-past/merged export; a plain causal-lm export only works in 'full sequence' mode. | |
| template | COMBO | chatml (Qwen2.5) | Prompt wrapper. chatml: Qwen2.5-Instruct (<|im_start|>user ...). gemma3: Gemma 3 instruct (<start_of_turn>user ..., prepends BOS). 'none': raw completion (GPT-2). The template also sets the default stop/BOS ids. |
| prompt | STRING | What is OpenCV? | User prompt, wrapped by 'template'. |
| max_new_tokens | INT | 321–2048 | Maximum tokens to generate. CPU inference is roughly a few tokens/second for a 0.5B model - keep this small. |
| decode_mode | COMBO | kv-cache: fast, one token per step through the graph, needs a causal-lm-with-past / merged export. full sequence: re-feeds the whole growing sequence every step (slow, but works with plain exports). | |
| lm_fileopt | COMBO | auto (from model folder) | Override which .onnx in the model folder is the language model. 'auto' picks the merged/only decoder export - set this only when the folder holds several and auto cannot tell. |
| tokenizeropt | COMBO | auto (from model folder) | Override the tokenizer folder. 'auto' uses the model folder's own config.json + tokenizer.json (or the nearest parent holding them); set this for an export that ships no tokenizer of its own. |
| stop_token_idsopt | STRING | Comma-separated ids that end generation. Blank = the template default (Qwen2.5: 151645, 151643; Gemma3: 1, 106; GPT-2: 50256 suggested). | |
| bos_token_idopt | INT | -1-1–1000000 | Token prepended to the encoded prompt. -1 = the template default (Gemma3 prepends 2; the others none). |
| engineopt | COMBO | new graph | DNN engine for cv2.dnn.readNetFromONNX. LLM ONNX graphs need the OpenCV 5 GRAPH engine; the classic 4.x engine cannot parse them. OpenCV 5.1 merged the two into one engine, so 'auto' is the graph engine there and this default follows the build. |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| response | STRING | The generated text only (prompt stripped, stop tokens removed). |
| full_text | STRING | Prompt + response, detokenized. |
| new_tokens | INT | Number of tokens generated (before stop-id removal). |