JLC CaptionForge Pipeline Planner
Where a CaptionForge run actually gets planned
- Input - single image
- single_image
- pipeline_plan
- pipeline_plan_json
CaptionForge's real workflow is batch captioning a folder of images for LoRA training, and that workflow is run by JLC CaptionForge Pipeline Planner. It's the front door: you point it at an image folder (or a single image), tell it how many witness captions to generate per image and from which engines, pick the distiller and validator models, set your LoRA trigger word, and it produces a pipeline_plan that every downstream CaptionForge node reads. Change the plan, and the whole run changes with it - that's the point of having one planning node instead of forty widgets scattered across the graph.
The output layout is worth knowing before you run anything: final TXT sidecars get written beside each source image, run artifacts (JSONL audit trails, run config, output paths) land in a run-specific working folder under your output root, and direct IMAGE inputs get copied into an opt_images/ folder so the validator can find them.
How it works
The planner emits a CAPTIONFORGE_PIPELINE_PLAN (also available as pipeline_plan_json, a string) that caption nodes and the capstone consume through their pipeline_plan pins. When connected, it overrides the standalone sampling settings on those nodes - shared seed schedules, temperature/top-p/top-k schedules, max image size, max new tokens. It does recursive folder traversal and filename glob filtering, so *.png over a nested directory is a normal day at the office.
The Planner - enabled toggle is a neat escape hatch: disable it and the node just passes through your optional IMAGE with an empty plan, so downstream nodes run in standalone mode without you having to bypass widgets by hand.
Inputs that matter
- Input - image path - image file or folder root; doubles as the validator's image root in planned runs.
- Input - recursive / Input - filename glob - how deep you go and which files count.
- Output - folder / Output - run name - where artifacts go (default
/tmp/ComfyUI/output/CaptionForge) and the prefix on every JSONL. - Caption - Joy / Qwen / Ollama runs/image - witness counts per image. The dropdowns cap at 5 on purpose, "to prevent accidental giant runs" per the author. Three witnesses at 5 runs each is 15 raw captions per image to distill - think before you max it out.
- Caption - base seed / seed mode - fixed, increment, decrement, or random across the run.
- Caption - temperature / top p / top k schedule - comma-separated schedules; the last value repeats if you run out.
- Caption - max new tokens - the shared token budget all witnesses get.
- Distiller - model / Validator - model - the Ollama tags for Pass B and Pass C (dropdowns from
config/captionforge_ollama_models.json, with Custom for arbitrary tags). - Final - caption style -
narrative,comma, orboth. - LoRA - trigger word / LoRA - user caption anchor - threaded through distiller and validator.
Outputs: single_image (your IMAGE passthrough), pipeline_plan, and pipeline_plan_json.
Install
Same pack install as everything else:
git clone https://github.com/Damkohler/CaptionForge.git ComfyUI/custom_nodes/CaptionForge
Restart ComfyUI. The planner itself loads no models, but the run it plans needs Ollama up with the distiller/validator tags pulled (ollama pull mistral-small:24b, ollama pull gemma4:26b).
Common issues
This node is where runs get misconfigured, so read your settings like a checklist: witness counts, seed mode, output folder, trigger word. If a caption node seems to ignore your standalone settings, check that the planner's overrides aren't stomping them - that's by design. And if a run produces nothing, the usual cause is a typo in Input - image path or a glob that matches zero files. Start with one image and low witness counts before you point it at a 10,000-image dataset, because the pack is unapologetically heavy.
Inputs (45)
| Name | Type | Default | Description |
|---|---|---|---|
| Planner - enabled | BOOLEAN | true | Enable CaptionForge Pipeline Planner mode. If disabled, this node passes through the optional IMAGE and emits an empty/falsy plan so downstream nodes can run in standalone mode without GUI bypassing. |
| Input - image path | STRING | Image file or image folder/root for ordinary folder/file workflows. This is also used as the validator image root. For quick single-image workflows, connect IMAGE to the optional Input - single image socket. | |
| Input - recursive | BOOLEAN | true | Whether captioning nodes should recurse when Input - image path is a folder. |
| Input - filename glob | STRING | * | Filename glob for folder captioning, e.g. *.png, *.jpg, or *. |
| Output - folder | STRING | /tmp/ComfyUI/output/CaptionForge | Output root folder. CaptionForge creates a run-specific working directory inside this folder for JSON/JSONL/audit artifacts. Final TXT sidecars are written beside their resolved source images. |
| Output - run name | STRING | captionforge_run | Run-root used to name config JSON, output path JSON, caption JSONL, distiller JSONL, prompt JSONLs, validator JSONL, and final JSONL. |
| Output - overwrite outputs | BOOLEAN | true | Overwrite generated run artifacts during capstone execution. |
| LoRA - trigger word | STRING | Optional shared LoRA trigger token/string preserved through the pipeline. | |
| LoRA - user caption anchor | STRING | Optional user style/identity anchor passed to distiller and validator. | |
| Caption - Joy runs/image | COMBO | 2 | Joy Caption runs per image. Set to Disabled to omit Joy from this run. Dropdown is capped at 5 to prevent accidental giant runs. |
| Caption - Qwen runs/image | COMBO | 2 | Qwen Caption runs per image. Set to Disabled to omit Qwen from this run. Dropdown is capped at 5 to prevent accidental giant runs. |
| Caption - Ollama runs/image | COMBO | Disabled | Ollama Caption runs per image for each connected JLC CaptionForge Ollama Caption node. The actual Ollama model tag is selected in each Ollama Caption node. Connecting multiple Ollama Caption nodes multiplies runtime, memory pressure, and raw-caption count. |
| Caption - base seed | INT | -1-1–4294967295 | Base seed for caption generation. -1 means unseeded when supported. |
| Caption - seed mode | COMBO | fixed | 4 options: fixed, increment, decrement, random |
| Caption - temperature schedule | STRING | 0.90 | Comma-separated caption temperature schedule; final value repeats if needed. |
| Caption - top p schedule | STRING | 0.60 | — |
| Caption - top k schedule | STRING | 80 | — |
| Caption - max image size | INT | 10240–4096 | — |
| Caption - max new tokens | INT | 600016–12000 | — |
| Distiller - model | COMBO | mistral-small:24b | Concrete Ollama text model tag for the distiller. Use custom to enter any other installed Ollama text model. |
| Distiller - custom Ollama model | STRING | Used only when Distiller - model is custom, e.g. my-model:latest. | |
| Distiller - base seed | INT | -1-1–4294967295 | Base seed for the distiller. -1 means omit seed. |
| Distiller - seed mode | COMBO | fixed | 4 options: fixed, increment, decrement, random |
| Distiller - strategy | COMBO | single_pass | 2 options: single_pass, by_model_then_global |
| Distiller - max caption chars for LLM | INT | 15360–12000 | — |
| Distiller - num predict | INT | 309664–12000 | — |
| Distiller - temperature | FLOAT | 0.240–2 | — |
| Distiller - top p | FLOAT | 0.900–1 | — |
| Distiller - top k | INT | 600–500 | — |
| Distiller - write prompt JSONL | BOOLEAN | false | — |
| Distiller - preserve raw response | BOOLEAN | false | — |
| Validator - model | COMBO | gemma4:26b | Concrete Ollama vision model tag for the image-aware validator. Use custom to enter any other installed Ollama vision model. |
| Validator - custom Ollama model | STRING | Used only when Validator - model is custom, e.g. gemma4:e4b or another installed VLM tag. | |
| Validator - base seed | INT | -1-1–4294967295 | Base seed for the VLM validator. -1 means omit seed. |
| Validator - seed mode | COMBO | fixed | 4 options: fixed, increment, decrement, random |
| Validator - num predict | INT | 220064–12000 | — |
| Validator - temperature | FLOAT | 0.000–2 | — |
| Validator - top p | FLOAT | 0.920–1 | — |
| Validator - top k | INT | 800–500 | — |
| Validator - write prompt JSONL | BOOLEAN | false | — |
| Validator - preserve raw VLM response | BOOLEAN | false | — |
| Final - caption style | COMBO | narrative | 3 options: narrative, comma, both |
| Final - write TXT sidecars | BOOLEAN | true | Final TXT sidecars use the associated image filename stem and .txt extension. |
| Final - write JSONL | BOOLEAN | true | — |
| Input - single imageopt | IMAGE | Optional IMAGE passthrough for quick single-image workflows. The planner does not process this image; it simply returns it as output. |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| single_image | IMAGE | — |
| pipeline_plan | CAPTIONFORGE_PIPELINE_PLAN | — |
| pipeline_plan_json | STRING | — |