Add Label (Swwan)
Burn text onto a whole batch without opening Photoshop
- image
- IMAGE
What it's actually for
You have 40 frames in a batch and you need to see which one is which. Or you're building a LoRA dataset and want the caption sitting on the image so you can eyeball whether the tag matches the picture. Add Label staples a strip of text onto every image in a batch - above, below, left, right, or straight on top - and hands you one batch back. It's the fastest way to make a contact sheet in ComfyUI without leaving the graph.
Worth knowing where this thing came from: it's KJNodes' AddLabel, re-registered under a Swwan node ID. The pack does that a lot - it ships its own copy of several well-known image/mask utilities under Swwan* names so the image-processing half of your graph doesn't need KJNodes, LayerStyle or 孤海 installed at all. Same widgets, same behaviour, different node ID.
How it works
The node builds a second image in PIL - a painted label canvas - writes your text into it with a TrueType font, then glues it to the input with torch.cat. For up and down that's a concat along the height axis; for left and right it rotates the label strip 90° and concats along the width axis. overlay is the odd one out: it draws the text directly onto a copy of your image and returns that, so the output is still the original resolution.
Text is word-wrapped to the image width (it measures each word and breaks lines), and the label height is either the height you set or, if you set height to -1, just enough rows to fit the wrapped lines plus an 8px margin. That -1 is the one people forget and then wonder why half their caption is clipped.
The inputs that matter
image and text are the obvious two. Then direction - up, down, left, right, overlay - and height, where -1 means auto-size. font_size is exactly what it says. font_color and label_color are PIL colour strings, so white, black, or a hex like #ffcc00 all work.
font is a dropdown built from the pack's own fonts folder, registered as swwan_fonts. There's a stale line in the node's description telling you it loads fonts from the KJNodes folder - ignore that, the code reads ComfyUI/custom_nodes/ComfyUI_Swwan/fonts. Two faces ship with it; drop your own .ttf/.otf in there and it shows up after a ComfyUI restart.
The optional caption input is a forceInput string, and it's the interesting one: it takes a list of strings, one per image in the batch, and asserts the length matches. Wire a list of captions in and each frame gets its own text instead of the same text on all of them. Get the count wrong and it fails loudly, which is honestly the right behaviour.
Output is a single IMAGE. Send it to Save Image, or a preview node, or straight into your dataset folder.
Install
ComfyUI Manager, search ComfyUI Swwan, or by hand:
cd ComfyUI/custom_nodes
git clone https://github.com/aining2022/ComfyUI_Swwan
cd ComfyUI_Swwan
python -m pip install -r requirements.txt # numpy, Pillow, opencv-python, scipy, scikit-image
That python has to be the interpreter you launch ComfyUI with - same rule for the Windows portable build, where you'd use .\python_embeded\python.exe. Then restart ComfyUI and hard-refresh the browser. No models, no downloads, and the pack won't touch your torch/CUDA install.
Where people get burned
The node changes your image height. The author says so in the description and it's still the number one surprise: label an image and suddenly everything downstream is 48px taller, so an img2img pass or a latent with baked-in dimensions is now mismatched. If you only want text in the preview, use overlay and keep the geometry intact.
Second: the font dropdown is populated at startup. Copying a new font in and refreshing the browser does nothing - restart.
Third, batch behaviour. Without caption, every image gets the same text. That's usually what you want for a contact sheet and almost never what you want for a dataset - wire per-image strings into caption there. And the node is pure PIL plus one concat per frame, so a 200-frame batch at 4K is a slow, memory-hungry loop; label a sample of the batch, not all of it, when you're just checking.
Inputs (11)
| Name | Type | Default | Description |
|---|---|---|---|
| image | IMAGE | — | |
| text_x | INT | 100–4096 | — |
| text_y | INT | 20–4096 | — |
| height | INT | 48-1–4096 | — |
| font_size | INT | 320–4096 | — |
| font_color | STRING | white | — |
| label_color | STRING | black | — |
| font | COMBO | 3 options: FreeMono.ttf, FreeMonoBoldOblique.otf, TTNorms-Black.otf | |
| text | STRING | Text | — |
| direction | COMBO | up | 5 options: up, down, left, right, overlay |
| captionopt | STRING | — |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| IMAGE | IMAGE | — |