HPSv3++ Caption
The write-the-prompt-for-me node, ++ edition
- model
- images
- captions
Sometimes you don't want the prompt you typed, you want the prompt that describes the image you got. HPSv3++ Caption reads a picture and writes a short description of it as a string you can wire straight into the rest of the graph.
What it is
The HPSv3++ model is a reward model built on a vision-language backbone - which is why it can hand you both a preference score and a caption. This node exposes the caption side. It's the ++ series version of HPSv3 Caption, and it's the captioner in the pack's headline example workflow, paired with HPSv3++ Score in a loop: the same image goes to both, and the generated description becomes the prompt the image is scored against.
That loop is the reason the node exists in this pack's story. Normally when you score an image you need to hand the scorer the prompt that was supposed to make it. If you're scoring somebody else's image, or an old render whose prompt you've lost, a captioner is how you reconstruct something close enough to evaluate against.
- LoRA dataset captioning. Get the first draft automatically, then edit. Auto-caption is a first pass, not a finished caption file.
- Prompt seeds. Not recovery - the README says plainly it doesn't reproduce the original generation prompt - but a good starting point for an img2img round trip.
- Describing outputs at scale. Caption a batch, skim the text, find the ones that came out weird.
Inputs and outputs
model- anHPSV3PP_MODEL, which means HPSv3++ Model Loader. The plain HPSv3 loader produces a different type and won't connect; the two series are intentionally incompatible.images- any IMAGE, batches included. Each image gets its own caption.max_new_tokens- default 96, range 16 to 512. It's a ceiling on the generated length, so a densely detailed scene can still come back truncated. If your captions keep stopping short, this is the dial.
One output: captions, a STRING list in the same order as the input images. Feed it to a CLIP Text Encode, to a text display node if you actually want to read it, or to HPSv3++ Score's prompt.
The list is positional, and that's the thing to be careful about. Pass the same images to Caption and Score in the same order and everything lines up; shuffle one of them and you'll be scoring image A against the caption of image B, silently. Nothing errors - you just get numbers that mean nothing.
Honest expectations
It writes fluent, plausible descriptions. That fluency is exactly what makes it easy to over-trust.
The known weakness across captioning VLMs - JoyCaption's own release notes call it out for the whole class, not just itself - is attribution when there's more than one subject in frame: who's wearing the hat, who's holding the cup, who's sitting down. Expect that to go wrong on crowded images.
The second thing, and it matters most if you're building a dataset: a caption that reads well still might not regenerate the source image. So for a small LoRA set, read them. For a large one, spot-check by captioning ten random files and reading those before you commit hours of training to a set an automated description flattened. People building serious datasets usually run a natural-language describer and a booru-style tagger and combine the outputs, rather than picking one.
Install
Manager: Manager → search type Node Pack → search ComfyUI-HPSv3 → confirm the repo is Stella2211/ComfyUI-HPSv3 → install → restart. That also installs transformers 5.17.x and bitsandbytes for you, and runs the pack's install.py step. Manual:
cd ComfyUI/custom_nodes
git clone https://github.com/Stella2211/ComfyUI-HPSv3.git
cd ComfyUI-HPSv3
uv pip install --python <ComfyUI Python> -r requirements.txt
uv run --no-project --python <ComfyUI Python> python install.py
The last command is only needed for manual installs, and don't clone over a Manager install.
Hardware: NVIDIA CUDA GPU with BF16 support, 12 GB VRAM as a guideline with 8 GB untested, Python 3.12+, and a driver that matches ComfyUI's PyTorch/CUDA. CPU, AMD and Apple GPUs are not supported. The model is around 6.5 GB and auto-downloads to ComfyUI/models/hpsv3pp/HPSv3-PlusPlus-bnb-NF4/ on the first run.
Quick sanity check after installing: restart, search HPSv3++ and confirm Model Loader, Score and Caption all show up. They're missing only if the extension failed to import - check the ComfyUI startup log.
Common issues
It looks hung on the first run. That's a 6.5 GB download. Progress prints to the terminal, not the browser. Cancelling is safe; partial data is kept and the download resumes next run.
Captions are empty or nonsense. Repair the extension's dependencies in Manager and restart, and leave PyTorch/torchvision alone - replacing them breaks more than it fixes here.
Captions end mid-sentence. Raise max_new_tokens.
Model socket refuses to connect. Wrong series. Use the ++ loader with the ++ captioner and scorer.
GPU memory errors. The reward/caption model holds VRAM while it runs. Caption in a separate, lighter workflow if you're tight, and close anything else using the card.
Inputs (3)
| Name | Type | Default | Description |
|---|---|---|---|
| model | HPSV3PP_MODEL | — | |
| images | IMAGE | — | |
| max_new_tokens | INT | 9616–512 | — |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| captions | STRING | — |