Imgutils WD14 Tagger
The tagger your anime training data actually wants
- image
- tags
- json
When people say "tagger" in the anime LoRA world, they mean WD14. It's the SmilingWolf family of taggers trained on Danbooru data, and it's the default captioner for the Illustrious/Pony/NoobAI lineage because it speaks exactly the Danbooru tag vocabulary those models were trained on - no natural-language descriptions, no BLIP-style rambling, just 1girl, long hair, blue eyes, school uniform in the order and format the model expects. Imgutils WD14 Tagger is that, as a ComfyUI node, part of the xiaden/comfyui-imgutils pack wrapping the deepghs/imgutils library. The library supports the different WD14 backbones (swinv2, convnext, moat, vit); this node wires you to that machinery with the sensible defaults.
Unlike the flat-output taggers in this pack (MLDanbooru, PixAI), WD14 splits its output into rating / general / character sections, and it gives each section its own confidence threshold. That split matters: character tags are the ones you care most about getting right, and they deserve a different bar than the general tags.
Inputs:
general_threshold(default 0.35) - confidence cutoff for general tags. The tooltip says it plainly: lower = more tags. 0.35 is chatty, which is the right default for training data (you'll filter later); raise it to ~0.5 for prompt-ready output.character_threshold(default 0.85) - cutoff for character tags. High by default because false-positive character tags are worse than missing ones - a wrong character name in your training caption teaches the model the wrong identity.drop_overlap(default off) - removes redundant tags.
Outputs:
tags- a comma-separated STRING with rating first, then general and character tags, each sorted by confidence. Ready to paste.json- the full{"rating": ..., "general": ..., "character": ...}dict with scores, for pipelines that need the numbers.
Why you'd reach for it
Captioning for a character LoRA on any Danbooru-tag base - this is the tool the lora-training essay points to explicitly for that job. The split output is the feature, honestly: being able to set a high character bar means your character tags stay clean while your general tags stay comprehensive, and that separation is what a good caption pipeline is built on.
Install & gotchas
cd ComfyUI/custom_nodes/
git clone https://github.com/xiaden/comfyui-imgutils.git
cd comfyui-imgutils
pip install -r requirements.txt
Restart ComfyUI. Requires ComfyUI 0.25.0+ (V3 node API pack). WD14's weights download from HuggingFace Hub on first use and cache to ~/.cache/huggingface/hub/ - first run is slow, offline fails, and a second tagger in the same pack may download its own separate weights.
Practical notes from actually captioning with this: drop_overlap can eat legitimate tags ("blue hair" vanishing because "blue hair long hair" got merged), so leave it off until you've seen a few outputs. And remember the character-tag trade - 0.85 is a good default, but if your reference images are stylized or partial shots, genuine character tags will sit below it; drop it to 0.7 and eyeball the JSON before you train on a big set.
Inputs (4)
| Name | Type | Default | Description |
|---|---|---|---|
| image | IMAGE | Input image to tag. | |
| general_threshold | FLOAT | 0.350–1 | Confidence threshold for general tags. Lower = more tags. |
| character_threshold | FLOAT | 0.850–1 | Confidence threshold for character tags. Higher = fewer false positives. |
| drop_overlap | BOOLEAN | false | Remove overlapping/redundant tags. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| tags | STRING | — |
| json | STRING | — |