Nodes/ComfyUI-HPSv3/HPSv3 Caption
ComfyUI Node

HPSv3 Caption

It writes you a prompt from a picture, and it isn't cheating

By Stella2211·Created 22 days ago·Updated 3 days ago· 2
HPSv3 Caption
  • model
  • images
  • captions
◄max_new_tokens96►

You have an image and you want the prompt back. HPSv3 Caption doesn't do that - and the author is upfront about it. What it does is write you a new short description of what's in the picture, which is often more useful than the original prompt anyway.

What it is

The HPSv3 checkpoints aren't just a number machine. They carry a tokenizer and a processor, which is the tell that there's a vision-language model under the hood - and this node exposes that side of it. Point it at an image, get text out. It's the same job as the captioner nodes people already run (JoyCaption, Florence-2, WD14 taggers), just packaged with the reward model you installed for scoring.

Where you'd actually reach for it:

  • Building caption files for a LoRA set. Generate the first draft, then fix it. That's the realistic workflow; nobody hand-writes 800 captions and nobody should trust an auto-captioner unread either.
  • Seeding img2img or image-to-video from an existing picture. Drop a photo in, get a description that's roughly the vocabulary your model understands, iterate from there.
  • The feedback loop this pack is built for: Caption into Score. The README's own example passes the same image to both nodes and wires Caption's output into Score's prompt, which evaluates the image against its own generated description. Handy when you don't have the prompt that made an image.
  • Vibe checks on a batch. Caption fifty outputs and skim the text; you'll spot the malformed ones faster than looking at the pictures.

Inputs and outputs

  • model - an HPSV3_MODEL from HPSv3 Model Loader. The ++ loader won't fit here; the two series are separate types on purpose.
  • images - anything IMAGE. Batches are fine and get captioned one by one.
  • max_new_tokens - default 96, adjustable from 16 to 512. This is the cap on how long the generated description can be, not a target. Bump it for cluttered scenes where the caption keeps ending mid-thought; drop it if you want short captions and faster runs.

The single output is captions: a STRING list, one entry per input image, in the order you fed them. Wire it into a CLIP Text Encode for img2img, or into HPSv3 Score's prompt socket. If you send multiple images through both nodes, pass the same images to each and keep the order identical - the list pairing is positional, and a mismatch here is the kind of bug that produces plausible-looking wrong results rather than an error.

To actually read what it produced, hang a text display node off the output. The pack's examples do this with ComfyUI's standard preview-as-text node, and it's the difference between "I think it wrote something" and knowing.

What it will and won't do

It will write a competent short description in natural language. It will not recover the generation prompt - the README says this explicitly, and it's not a limitation the author is planning to fix, it's a property of what the task is. A picture does not contain its prompt; the captioner is guessing at a description, and it's guessing from what it learned about images.

Two failure modes to expect, both shared with every other captioner in this space:

Multi-subject confusion. Put two people in a frame and the model will happily swap who's wearing what and who's doing what. This is the well-known weak spot of captioning VLMs, and it's worth knowing before you auto-caption a dataset full of group shots.

Captions that read plausibly but wouldn't regenerate the image. The output is fluent enough to feel authoritative. Hand-check a sample before you trust a whole set - or do what most people end up doing in practice: run a describer like this alongside a tagger, and concatenate.

Install

Manager: open Manager, set search type to Node Pack, search ComfyUI-HPSv3, verify the repo is Stella2211/ComfyUI-HPSv3, install, restart. Manager also pulls the dependencies (transformers 5.17.x, bitsandbytes) and runs the pack's setup step. Manual:

cd ComfyUI/custom_nodes
git clone https://github.com/Stella2211/ComfyUI-HPSv3.git
cd ComfyUI-HPSv3
uv pip install --python <ComfyUI Python> -r requirements.txt
uv run --no-project --python <ComfyUI Python> python install.py

Requirements: NVIDIA CUDA with BF16, 12 GB VRAM as a guideline (8 GB untested), Python 3.12+. No CPU, AMD or Apple GPU support. The 5.9 GB model auto-downloads on first run into ComfyUI/models/hpsv3/HPSv3-bnb-NF4/.

Common issues

First run crawls. The model is downloading. Watch the terminal.

transformers or bitsandbytes errors on load. Repair the extension's dependencies in Manager and restart. Don't swap PyTorch or torchvision; the pack uses ComfyUI's builds.

Captions cut off mid-sentence. You hit max_new_tokens. Raise it.

Type error on the model socket. You've connected an HPSv3++ loader. Match the series.

CUDA out of memory. Close other GPU workloads, or caption in a separate workflow from the one generating. Big video models and this node do not want to be neighbours.

CategoryHPSv3

Inputs (3)

NameTypeDefaultDescription
modelHPSV3_MODEL—
imagesIMAGE—
max_new_tokensINT9616–512—

Outputs (1)

NameTypeDescription
captionsSTRING—