Nodes/ComfyUI CV/CV DNN Tokenizer
ComfyUI Node

CV DNN Tokenizer

See what your prompt turns into before the model does

By bmad4ever·Created 4 months ago·Updated 15 days ago· 1
CV DNN Tokenizer
  • token_ids
  • token_ids
  • text
  • count
◄mode▾►
◄tokenizer▾►
◄textWhat is OpenCV?►

Why you'd reach for this

OpenCV 5 grew a tokenizer - cv2.dnn.Tokenizer - and this pack exposes it as a node. Drop a model folder in, get token ids out; feed ids back in, get text. Nothing generates here. It's the equivalent of holding your prompt up to the light before the model sees it.

Two reasons that's more useful than it sounds. First, token counts are the budget: the LLM/VLM/Seq2Seq nodes in this pack generate in token space, and knowing that your system prompt is eating 300 of the context tells you why the answer is truncated. Second, it's the only way to check what a prompt actually tokenizes into - the gap between what you typed and what the model reads is where prompt debugging lives. The KB's prompt-engineering notes make the same point about punctuation and weight syntax being passed through literally; seeing the ids is how you stop guessing.

How it works

encode runs tokenizer.encode(text) and packs the ids into a (1, N) int32 array. decode takes (N,) or (1, N) integer ids and returns the string. It calls into OpenCV's implementation, so it's BPE for the gpt2/gpt4/qwen2.5 families and SentencePiece for Gemma - whichever the model folder declares.

The tokenizer dropdown lists every folder under ComfyUI/models/llm that has both an OpenCV-format config.json and the family's tokenizer.json. Both parts matter: this is not a stock HuggingFace folder. The OpenCV-format config and tokenizer files come from opencv_extra (testdata/dnn/llm/<family>/), and the repo's model_sources.txt spells out which files go with which model (Qwen2.5-0.5B-Instruct needs them alongside an ONNX export; PaliGemma2 needs the gemma2 SentencePiece tokenizer). Paths are relative to the model root, so a workflow that names opencv/qwen2.5-0.5b-instruct stays portable between machines.

Inputs and outputs that matter

  • mode - encode (text -> ids) or decode (ids -> text). Note decode without token_ids wired raises rather than returning an empty string, which is the behaviour you want.
  • tokenizer - the model folder name under models/llm, not a file.
  • text - the prompt to encode. It comes back on the text output unchanged in encode mode, which is handy: the node is a pass-through plus a tap.
  • token_ids - the (1, N) int32 ids in encode mode (an empty (1, 0) when decoding), and the thing you wire in to decode.
  • count - tokens encoded or decoded. This is the one to actually watch. Wire it into a text/note node or a dashboard and you have a live prompt-length meter.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"

Python ≥ 3.12, ComfyUI on the V3 node API, or search ComfyUI CV in ComfyUI Manager. Then build the model folder by hand:

ComfyUI/models/llm/opencv/qwen2.5-0.5b-instruct/
  model.onnx                 # from onnx-community/Qwen2.5-0.5B-Instruct
  config.json                # OpenCV-format, from opencv_extra testdata/dnn/llm/
  tokenizer.json             # the family's, from the same place

If the dropdown is empty, it's almost always the second or third file that's missing - the node filters the folder list on both.

Common issues

  • Empty dropdown. No folder under models/llm has both required files. Check the case of the path too; the combo lists folders, not files.
  • Decode returns mojibake. You're decoding with a different family's tokenizer than the one that encoded the ids. BPE and SentencePiece do not share vocabularies.
  • Ids that look wrong for a "simple" prompt. BPE splits on things you didn't type deliberately - the same effect the KB documents for underscore-heavy negatives in CLIP. Encode it, look at count, and adjust.
  • You wanted generation, not tokenization. Use CV LLM Generate / CV VLM Generate / CV Seq2Seq Generate; they tokenize internally. This node is the microscope, not the engine.
  • The contrib build. Every node in the pack needs opencv-contrib-python-headless; installing a non-contrib wheel over it empties the contrib submodules and contrib nodes vanish. tools/repair_opencv_contrib.py --check then --apply if that happened.
Categoryimage/CV/dnn

Inputs (4)

NameTypeDefaultDescription
modeCOMBOencode: 'text' -> token_ids. decode: the wired 'token_ids' -> text.
tokenizerCOMBOModel folder under models/llm holding the family's config.json + tokenizer.json (e.g. opencv/qwen2.5-0.5b-instruct). Relative to the model root, so the workflow stays portable.
textoptSTRINGWhat is OpenCV?Text to encode (encode mode).
token_idsoptNPARRAY(N,) or (1, N) integer token ids to decode (decode mode).

Outputs (3)

NameTypeDescription
token_idsNPARRAY(1, N) int32 token ids (encode mode; empty (1, 0) in decode mode).
textSTRINGDecoded text (decode mode; the input text echoed in encode mode).
countINTNumber of tokens encoded / decoded.