JLC CaptionForge Joy Caption
Joy Caption is the witness you'll reach for first
- image
- pipeline_plan
- template_options
- image_out
- pipeline_plan_out
- template_options_out
- caption
- resolved_prompt
JoyCaption is the LoRA crowd's default captioner for a reason: it's the uncensored vision-language model built for training captions, and it's been the community favorite for years. This node is that model wrapped for ComfyUI, and inside CaptionForge it plays the role of the first "witness" - the Pass A caption engine whose account gets weighed against the others later. It's also a perfectly good standalone JoyCaption if you never touch the rest of the pack.
Here's the thing you should know first: CaptionForge is a brand-new, still-labeled-v0.1.x project (2026, MIT, by J. L. Córdova), so googling "CaptionForge" will mostly turn up an unrelated folder-tracking tool with the same name. Make sure you're on Damkohler/CaptionForge before you clone.
How it works
Under the hood it's a Python/Hugging Face engine running the JoyCaption/LLaVA-family model. Models are expected in ComfyUI/models/LLM/JLC_JoyCaption/, and if the folder's empty the node will download the weights from Hugging Face on first use - that's a multi-GB pull, not a typo on your part. The download_probe_only toggle, sitting at the very bottom of the widget list by design, fetches just the lightweight metadata and returns a status message instead of captioning. Use it to check your wiring before committing to a full download.
The node has two personalities. With nothing connected it's a standalone captioner. Plug a pipeline_plan in and it switches into Pass A evidence mode: the Pipeline Planner takes over image routing, per-run seeds, sampling schedules, and shared output paths, and your caption gets appended to the run's JSONL evidence file. Most of the sampling widgets stop mattering in that mode - the planner overrides them.
Inputs and outputs that matter
The few you'll actually touch:
- model - defaults to
llama-joycaption-beta-one-hf-llava. There's a model registry, so the dropdown may grow as models get added. - memory_mode -
Balanced (8-bit)is the default and the author's recommendation for 16 GB VRAM. It's bitsandbytes load-time quantization. If you can't load the 8-bit version,Defaultis your fallback. - caption_type / caption_length - the structured template path. Joy has thirteen caption styles (Descriptive, MidJourney, Danbooru tag list, Art Critic…) and a length target from "very short" to explicit token counts.
- custom_prompt_mode - overrides the template path when enabled;
custom_promptwins overprompt_presetif non-empty. - forbidden_phrases / replace_pairs - a cheap cleanup filter. One phrase per line to drop captions containing it, or
old=>newlines for replacements.
Outputs: caption (the text you actually want), resolved_prompt (what was actually sent to the model - handy for debugging), plus passthroughs image_out, pipeline_plan_out, and template_options_out so you can chain nodes cleanly. Wire caption into a text display or into the capstone's JSONL feed.
Install
One install serves the whole pack. ComfyUI Manager, search "CaptionForge", or:
git clone https://github.com/Damkohler/CaptionForge.git ComfyUI/custom_nodes/CaptionForge
Restart ComfyUI, then make sure the Python deps are in your ComfyUI environment:
cd ComfyUI/custom_nodes/CaptionForge
pip install -e .
The repo ships no weights - Joy downloads itself into ComfyUI/models/LLM/JLC_JoyCaption/. Optional 8-bit support needs pip install bitsandbytes.
Common issues
First run downloads a big model, so give it time or use download_probe_only to test the graph first. On 16 GB systems keep memory_mode on Balanced (8-bit). And if you're running the full pipeline, know that Joy and Qwen are process-local Python models while the distillation stages live in Ollama - CaptionForge clears the Python models before handing work to Ollama, so don't be alarmed if the next image feels like a fresh load. That eviction dance is the pack trying to fit two model ecosystems in one GPU.
Inputs (23)
| Name | Type | Default | Description |
|---|---|---|---|
| model | COMBO | llama-joycaption-beta-one-hf-llava | JoyCaption/LLaVA-family model. Models are loaded from ComfyUI/models/LLM/JLC_JoyCaption/. Missing models may be downloaded automatically unless download_probe_only is enabled. |
| memory_mode | COMBO | Balanced (8-bit) | Joy model memory mode. Balanced (8-bit) uses bitsandbytes load-time quantization and is the recommended CaptionForge default for 16 GB VRAM systems. |
| keep_loaded | BOOLEAN | true | Keep the model cached after captioning for faster repeated runs. CaptionForge cache policy may still evict it when another caption model must load. |
| caption_template_mode | BOOLEAN | true | Use the structured CaptionForge template path: caption_type, caption_length, and optional Template Options from the template_options pin. If custom_prompt_mode is also enabled, custom_prompt_mode takes precedence. |
| caption_type | COMBO | JLC LoRA Literal | Caption template style used when caption_template_mode is active. |
| caption_length | COMBO | any | Target caption length used when caption_template_mode is active. |
| custom_prompt_mode | BOOLEAN | false | Use custom_prompt when non-empty, otherwise use prompt_preset. This overrides caption_template_mode when both toggles are enabled. |
| prompt_preset | COMBO | default_literal | Built-in prompt preset used only in custom_prompt_mode when custom_prompt is blank. |
| system_prompt | STRING | You are a helpful image-captioning assistant. Describe only what is visible in the image. Do not invent unseen context. | Joy/LLaVA system prompt. Kept next to custom_prompt because both control the instruction envelope. Pipeline Planner does not currently override this. |
| custom_prompt | STRING | Custom prompt used only when custom_prompt_mode is enabled. Overrides prompt_preset when non-empty. | |
| max_new_tokens | INT | 38416–4096 | Standalone token budget. When a Pipeline Planner is connected, this is overridden by the Planner's shared max_new_tokens. |
| temperature | FLOAT | 0.750–2 | Standalone sampling temperature. When a Pipeline Planner is connected, this is overridden by the Planner temperature schedule. |
| top_p | FLOAT | 0.900–1 | Standalone top-p sampling value. When a Pipeline Planner is connected, this is overridden by the Planner top-p schedule. |
| top_k | INT | 500–500 | Standalone top-k sampling limit. When a Pipeline Planner is connected, this is overridden by the Planner top-k schedule. |
| repetition_penalty | FLOAT | 1.001–2 | Penalty applied to repeated tokens. Kept with the core captioning parameters. This is not currently overridden by the Pipeline Planner. |
| max_size | INT | 10240–4096 | Maximum longest-side image size for standalone captioning. The image is resized in memory only. Pipeline Planner overrides this in planned runs. |
| forbidden_phrases | STRING | Optional cleanup filter: remove lines/captions containing any listed phrase, one per line. | |
| replace_pairs | STRING | Optional cleanup replacements, one per line: old=>new. | |
| download_probe_only | BOOLEAN | false | At the very bottom by design. Probe/download lightweight model metadata only, then return a status message without captioning. |
| imageopt | IMAGE | Image or batch of images to caption. The image is passed through unchanged for clean node-to-node pipeline chaining. | |
| pipeline_planopt | CAPTIONFORGE_PIPELINE_PLAN | Connect the CaptionForge Pipeline Planner output here. When connected, this node switches into Pass A evidence mode: Planner image routing, per-run seeds, sampling schedules, shared output paths, and internal JSONL evidence append. | |
| template_optionsopt | CAPTIONFORGE_EXTRA_OPTIONS | Connect the CaptionForge Template Options node here. Works in standalone and Pipeline modes. This is the only source for template modifiers and name input. | |
| seedopt | INT | Optional standalone seed input. Ignored when a Pipeline Planner supplies a seed schedule. |
Outputs (5)
| Name | Type | Description |
|---|---|---|
| image_out | IMAGE | — |
| pipeline_plan_out | CAPTIONFORGE_PIPELINE_PLAN | — |
| template_options_out | CAPTIONFORGE_EXTRA_OPTIONS | — |
| caption | STRING | — |
| resolved_prompt | STRING | — |