cv2.putText
Burn text into a frame — and the two defaults that make it invisible
- source
- org
- result
There is no font engine here, no layout, no text wrapping, no anti-aliased kerning with a nice typeface. cv2.putText draws ASCII with Hershey stroke fonts, straight into the pixels. That limitation is also its virtue: it's deterministic, dependency-free, and works on any image or batch without a single model or node that can break. Frame counters on a video comparison, labels on a contact sheet, coordinates on a debug render, a quick "v2 / denoise 0.35" tag on an A-B test - that's the job.
The mechanism, and the two traps
Under the hood it takes a canvas, a string, an origin point (the bottom-left corner of the text, which trips people coming from anything else), a font face, a scale factor, and a colour in OpenCV's scalar format.
fontScaledefaults to 0 and 0 means invisible. A scale of 0 multiplies the font's base size to nothing. If your node runs cleanly and the image is unchanged, this is why.1.0is normal-sized text on a 512–1024px frame;0.5for small labels;2.0+ for titles.coloris a string literal in BGR, e.g."(0, 255, 0)"for green,"(0, 0, 255)"for red,"(255, 255, 255)"for white. A bare number broadcasts to every component, so"255"is white and"0"is black. The default(0, 0, 0, 0)is transparent-black-on-nothing - with a 3-channel canvas the fourth component is simply ignored, and you get black text. Black text on a dark frame is the second thing people report as "nothing happened".
The reason BGR matters: ComfyUI holds images as RGB tensors, but the wrapper converts an IMAGE link to uint8 BGR before calling cv2, so the channel order you type is the channel order OpenCV sees. That's consistent, not confusing - just don't assume the first number is red.
Inputs and outputs
Required: source (the canvas - IMAGE, MASK or NPARRAY), text (a plain string), org (a CV_TUPLE, i.e. x, y typed in place or wired from a CV Tuple node), fontFace (a dropdown of the Hershey fonts, default FONT_HERSHEY_SIMPLEX - the plainest and most legible of them), fontScale, and color as a literal string.
Optional: thickness (int, default 1), lineType (default LINE_AA, which is what makes the edges soft rather than jagged - switch to the non-anti-aliased types only if you want that pixelated look), and bottomLeftOrigin (default false; flip it if you're feeding an image whose rows are already bottom-up, which is rare).
The output is result, and it echoes the input's format: an IMAGE in comes back as an IMAGE, a MASK as a MASK, an NPARRAY stays an NPARRAY. So it drops into a graph without an adapter node, which is not true of every wrapper in this pack.
One behaviour worth knowing: putText is in the pack's per-frame batch set. Hand it an IMAGE batch of a clip and every frame gets the same text, which is exactly right for a watermark and exactly wrong if you wanted a counter on one frame. For the latter, convert with a batch index or slice first.
Install
From ComfyUI CV (bmad4ever). ComfyUI Manager, search comfyui_cv, or:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"
Restart. Python ≥ 3.12, recent V3-API ComfyUI, no models - the Hershey fonts are built into OpenCV.
Common issues
Nothing appears. fontScale = 0, or the text is black on a black frame, or org is outside the image (remember it's the bottom-left corner; a y of 0 puts the text above the frame).
Accented characters come out as boxes or question marks. The Hershey fonts are for basic Latin. Anything past that - accented letters, CJK, emoji - will not render. For real typography, composit in an image editor or use a node that rasterises a font.
The text is on every frame of my video. That's the batch behaviour described above; the node applies to all frames by design.
Garbled output colour. Check the BGR order. (255, 0, 0) is blue in a canvas that was converted from RGB, not red.
The pack is missing nodes entirely. If the node isn't in the menu at all, it's usually an OpenCV wheel problem - a non-contrib opencv-python installed over a contrib one empties the contrib modules. tools/repair_opencv_contrib.py --check inside the pack folder, then --apply.
Inputs (9)
| Name | Type | Default | Description |
|---|---|---|---|
| source | COMFY_MATCHTYPE_V3 | The image output(s) echo this input's format. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| text | STRING | Text string to be drawn. | |
| org | CV_TUPLE | 0,0 | Bottom-left corner of the text string in the image. One value with 2 components (x, y) - it travels as a whole, so it cannot arrive half-connected. Wire it from 'CV Tuple' or type the components in place. |
| fontFace | COMBO | FONT_HERSHEY_SIMPLEX | Font type, see #HersheyFonts. |
| fontScale | FLOAT | 0.0000-1e+38–1e+38 | Font scale factor that is multiplied by the font-specific base size. |
| color | STRING | (0, 0, 0, 0) | Text color. cv2 Scalar as a literal, e.g. "(0, 255, 0)" (BGR) or "(0, 255, 0, 64)" (BGRA). A bare number broadcasts to every component, so "255" means (255, 255, 255, 255). Components past the target's channel count are ignored by OpenCV. |
| thicknessopt | INT | 1-2147483648–2147483647 | Thickness of the lines used to draw a text. Preset to the OpenCV default (1). |
| lineTypeopt | COMBO | LINE_AA | Line type. See #LineTypes |
| bottomLeftOriginopt | BOOLEAN | false | When true, the image data origin is at the bottom-left corner. Otherwise, it is at the top-left corner. Preset to the OpenCV default (False). |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| result | COMFY_MATCHTYPE_V3 | Echoes the 'source' input's format: an IMAGE link comes back as IMAGE, MASK as MASK, NPARRAY stays NPARRAY. |