CV DNN Tokenizer
See what your prompt turns into before the model does
- token_ids
- token_ids
- text
- count
Why you'd reach for this
OpenCV 5 grew a tokenizer - cv2.dnn.Tokenizer - and this pack exposes it as a node. Drop a model folder in, get token ids out; feed ids back in, get text. Nothing generates here. It's the equivalent of holding your prompt up to the light before the model sees it.
Two reasons that's more useful than it sounds. First, token counts are the budget: the LLM/VLM/Seq2Seq nodes in this pack generate in token space, and knowing that your system prompt is eating 300 of the context tells you why the answer is truncated. Second, it's the only way to check what a prompt actually tokenizes into - the gap between what you typed and what the model reads is where prompt debugging lives. The KB's prompt-engineering notes make the same point about punctuation and weight syntax being passed through literally; seeing the ids is how you stop guessing.
How it works
encode runs tokenizer.encode(text) and packs the ids into a (1, N) int32 array. decode takes (N,) or (1, N) integer ids and returns the string. It calls into OpenCV's implementation, so it's BPE for the gpt2/gpt4/qwen2.5 families and SentencePiece for Gemma - whichever the model folder declares.
The tokenizer dropdown lists every folder under ComfyUI/models/llm that has both an OpenCV-format config.json and the family's tokenizer.json. Both parts matter: this is not a stock HuggingFace folder. The OpenCV-format config and tokenizer files come from opencv_extra (testdata/dnn/llm/<family>/), and the repo's model_sources.txt spells out which files go with which model (Qwen2.5-0.5B-Instruct needs them alongside an ONNX export; PaliGemma2 needs the gemma2 SentencePiece tokenizer). Paths are relative to the model root, so a workflow that names opencv/qwen2.5-0.5b-instruct stays portable between machines.
Inputs and outputs that matter
- mode -
encode (text -> ids)ordecode (ids -> text). Note decode withouttoken_idswired raises rather than returning an empty string, which is the behaviour you want. - tokenizer - the model folder name under
models/llm, not a file. - text - the prompt to encode. It comes back on the
textoutput unchanged in encode mode, which is handy: the node is a pass-through plus a tap. - token_ids - the
(1, N)int32 ids in encode mode (an empty(1, 0)when decoding), and the thing you wire in to decode. - count - tokens encoded or decoded. This is the one to actually watch. Wire it into a text/note node or a dashboard and you have a live prompt-length meter.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"
Python ≥ 3.12, ComfyUI on the V3 node API, or search ComfyUI CV in ComfyUI Manager. Then build the model folder by hand:
ComfyUI/models/llm/opencv/qwen2.5-0.5b-instruct/
model.onnx # from onnx-community/Qwen2.5-0.5B-Instruct
config.json # OpenCV-format, from opencv_extra testdata/dnn/llm/
tokenizer.json # the family's, from the same place
If the dropdown is empty, it's almost always the second or third file that's missing - the node filters the folder list on both.
Common issues
- Empty dropdown. No folder under
models/llmhas both required files. Check the case of the path too; the combo lists folders, not files. - Decode returns mojibake. You're decoding with a different family's tokenizer than the one that encoded the ids. BPE and SentencePiece do not share vocabularies.
- Ids that look wrong for a "simple" prompt. BPE splits on things you didn't type deliberately - the same effect the KB documents for underscore-heavy negatives in CLIP. Encode it, look at
count, and adjust. - You wanted generation, not tokenization. Use
CV LLM Generate/CV VLM Generate/CV Seq2Seq Generate; they tokenize internally. This node is the microscope, not the engine. - The contrib build. Every node in the pack needs
opencv-contrib-python-headless; installing a non-contrib wheel over it empties the contrib submodules and contrib nodes vanish.tools/repair_opencv_contrib.py --checkthen--applyif that happened.
Inputs (4)
| Name | Type | Default | Description |
|---|---|---|---|
| mode | COMBO | encode: 'text' -> token_ids. decode: the wired 'token_ids' -> text. | |
| tokenizer | COMBO | Model folder under models/llm holding the family's config.json + tokenizer.json (e.g. opencv/qwen2.5-0.5b-instruct). Relative to the model root, so the workflow stays portable. | |
| textopt | STRING | What is OpenCV? | Text to encode (encode mode). |
| token_idsopt | NPARRAY | (N,) or (1, N) integer token ids to decode (decode mode). |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| token_ids | NPARRAY | (1, N) int32 token ids (encode mode; empty (1, 0) in decode mode). |
| text | STRING | Decoded text (decode mode; the input text echoed in encode mode). |
| count | INT | Number of tokens encoded / decoded. |