ComfyUI Node

PixAI Tagger

Image in, one comma-separated tag string out

By warwk·Created 2 days ago·Updated 2 days ago· 2
PixAI Tagger
  • images
  • all_tags
◄modelpixai-labs/pixai-tagger-v1.0►
◄threshold0.17►
◄character_threshold0.27►
◄replace_underscorefalse►
◄trailing_commafalse►
◄unload_modelfalse►
◄exclude_tags►

WD14 has been the default anime tagger for so long that "run the tagger" stopped meaning anything else. PixAI Tagger v1.0 is the newer thing people keep pointing at - a 486M-parameter model over 30,877 tags - and this node is its simple front door: images in, one comma-separated string out, shaped like the WD14 node most workflows already have.

What it actually does, and why you'd swap

Tags aren't descriptions. If you're captioning for Illustrious, NoobAI or Pony-lineage models, booru tags match how the base was trained, which is exactly why the standard practice for a tag model is a tagger and a describer side by side rather than one or the other. This is the tagger half.

What makes it worth a look over WD14-class taggers:

  • A style head. The vocabulary has 4,917 style labels, and the wider benchmark table on the model card shows most tagger families have no style output at all.
  • A bigger, newer vocabulary - 30,877 tags, with a knowledge cutoff of May 2026, so recent characters actually resolve.
  • PixAI's card puts it top on general micro-F1 in its own eight-model comparison of shared tags. Vendor benchmarks, one test set, no confidence intervals - read it as "credible, not settled".

The trade is speed: 1008×1008 on a 486M backbone. The card's H100 setup managed 48.3 images/s at batch 16 while every model it beat ran faster; on a consumer card you feel that gap per image.

How it works

The node builds a transformers pipeline straight from the Hugging Face repo with trust_remote_code=True - the model repo ships its own pipeline class, which is why the pack needs timm installed. Your image gets resized and padded to 1008×1008 (aspect ratio preserved, no stretching) and scored as multi-label over all six categories.

This node then takes three of them, general, character and style, filters each at its threshold, sorts what survives by confidence descending, escapes parentheses so the tags are prompt-safe, and joins everything with , . General tags come first, then characters, then style. 1girl, solo, long hair, looking at viewer is the shape you get.

The inputs you'll actually touch

  • threshold (0.17 default) - the cut for general and style tags. Lower means more tags. This is the one you tune.
  • character_threshold (0.27 default) - separate, and higher, because character names are high-confidence picks.
  • replace_underscore - the vocabulary stores long_hair; turn this on if the captions or prompts it feeds use spaces.
  • exclude_tags - comma-separated, case-insensitive. text, watermark, signature is the usual starter list.
  • unload_model - frees the tagger after the run. Worth flipping on if the sampler is about to want the VRAM back.

Output: all_tags, a STRING. One entry per image in your batch, so a batch of four gives you four strings - and anything downstream runs once per image, which is what you want for batch captioning and mildly surprising the first time you expected a single merged prompt.

Install

ComfyUI Manager, search ComfyUI PixAI Tagger (it's on the registry), install, restart. Or by hand:

cd ComfyUI/custom_nodes
git clone https://github.com/warwk/ComfyUI-PixAI-Tagger.git
pip install -r requirements.txt

That pulls transformers, timm, huggingface_hub, numpy and Pillow. The README is explicit that torch is not a dependency on purpose - don't install or replace it by hand to satisfy this node.

The weights aren't bundled. First run downloads pixai-labs/pixai-tagger-v1.0 from Hugging Face - about 1.95 GB of safetensors into your HF cache - and caches it. Second run is instant.

Where people get burned

It looks frozen the first time. There's no progress bar for a ~2 GB download; check your console or ~/.cache/huggingface before you assume it hung. Corporate networks and HF_HUB_OFFLINE setups fail here too, and trust_remote_code=True means the model repo's own code has to be allowed to run.

Your exclude list must match the output, not the vocabulary. Exclusion is applied after underscore replacement. If replace_underscore is on and you put long_hair in the list, nothing gets excluded - you needed long hair. Conversely, with it off, the spaced version is the one that misses.

The threshold is strict. A tag has to beat the value, not equal it, and the widget steps in 0.05.

No rating or copyright tags, ever. They're not filtered out, they're simply not in this node's output. If you wanted rating:s for NSFW routing or franchise tags for a dataset, that's the other node in the pack.

Anime illustrations only. PixAI's own limitations section says they haven't measured it on other image types, and anything newer than May 2026 may simply be absent.

If it crawls, that's the 1008² resolution, not a misconfiguration - and a third-party ONNX build of this model exists precisely because people wanted it faster.

Categoryimage

Inputs (8)

NameTypeDefaultDescription
imagesIMAGEThe input images to be tagged.
modelCOMBOpixai-labs/pixai-tagger-v1.0The model to use for tagging. Currently pixai-labs/pixai-tagger-v1.0 is the only available model.
thresholdFLOAT0.170–1Threshold for general tags. Lower values will include more tags, higher values will be more selective. Default is 0.17.
character_thresholdFLOAT0.270–1Threshold for character tags. Lower values will include more tags, higher values will be more selective. Default is 0.27.
replace_underscoreBOOLEANfalseIf enabled, underscores in tags will be replaced with spaces.
trailing_commaBOOLEANfalseIf enabled, a trailing comma will be added after the last tag.
unload_modelBOOLEANfalseIf enabled, the model will be unloaded after tagging.
exclude_tagsSTRINGA comma-separated list of tags to exclude from the results.

Outputs (1)

NameTypeDescription
all_tagsSTRING—