Nodes/CaptionForge/ JLC CaptionForge Pipeline Planner
ComfyUI Node

 JLC CaptionForge Pipeline Planner

Where a CaptionForge run actually gets planned

By Damkohler·Created 3 months ago·Updated 2 months ago· 1
 JLC CaptionForge Pipeline Planner
  • Input - single image
  • single_image
  • pipeline_plan
  • pipeline_plan_json
Planner - enabledtrue
Input - image path
Input - recursivetrue
Input - filename glob*
Output - folder/tmp/ComfyUI/output/CaptionForge
Output - run namecaptionforge_run
Output - overwrite outputstrue
LoRA - trigger word
LoRA - user caption anchor
Caption - Joy runs/image2
Caption - Qwen runs/image2
Caption - Ollama runs/imageDisabled
Caption - base seed-1
Caption - seed modefixed
Caption - temperature schedule0.90
Caption - top p schedule0.60
Caption - top k schedule80
Caption - max image size1024
Caption - max new tokens6000
Distiller - modelmistral-small:24b
Distiller - custom Ollama model
Distiller - base seed-1
Distiller - seed modefixed
Distiller - strategysingle_pass
Distiller - max caption chars for LLM1536
Distiller - num predict3096
Distiller - temperature0.24
Distiller - top p0.90
Distiller - top k60
Distiller - write prompt JSONLfalse
Distiller - preserve raw responsefalse
Validator - modelgemma4:26b
Validator - custom Ollama model
Validator - base seed-1
Validator - seed modefixed
Validator - num predict2200
Validator - temperature0.00
Validator - top p0.92
Validator - top k80
Validator - write prompt JSONLfalse
Validator - preserve raw VLM responsefalse
Final - caption stylenarrative
Final - write TXT sidecarstrue
Final - write JSONLtrue

CaptionForge's real workflow is batch captioning a folder of images for LoRA training, and that workflow is run by JLC CaptionForge Pipeline Planner. It's the front door: you point it at an image folder (or a single image), tell it how many witness captions to generate per image and from which engines, pick the distiller and validator models, set your LoRA trigger word, and it produces a pipeline_plan that every downstream CaptionForge node reads. Change the plan, and the whole run changes with it - that's the point of having one planning node instead of forty widgets scattered across the graph.

The output layout is worth knowing before you run anything: final TXT sidecars get written beside each source image, run artifacts (JSONL audit trails, run config, output paths) land in a run-specific working folder under your output root, and direct IMAGE inputs get copied into an opt_images/ folder so the validator can find them.

How it works

The planner emits a CAPTIONFORGE_PIPELINE_PLAN (also available as pipeline_plan_json, a string) that caption nodes and the capstone consume through their pipeline_plan pins. When connected, it overrides the standalone sampling settings on those nodes - shared seed schedules, temperature/top-p/top-k schedules, max image size, max new tokens. It does recursive folder traversal and filename glob filtering, so *.png over a nested directory is a normal day at the office.

The Planner - enabled toggle is a neat escape hatch: disable it and the node just passes through your optional IMAGE with an empty plan, so downstream nodes run in standalone mode without you having to bypass widgets by hand.

Inputs that matter

  • Input - image path - image file or folder root; doubles as the validator's image root in planned runs.
  • Input - recursive / Input - filename glob - how deep you go and which files count.
  • Output - folder / Output - run name - where artifacts go (default /tmp/ComfyUI/output/CaptionForge) and the prefix on every JSONL.
  • Caption - Joy / Qwen / Ollama runs/image - witness counts per image. The dropdowns cap at 5 on purpose, "to prevent accidental giant runs" per the author. Three witnesses at 5 runs each is 15 raw captions per image to distill - think before you max it out.
  • Caption - base seed / seed mode - fixed, increment, decrement, or random across the run.
  • Caption - temperature / top p / top k schedule - comma-separated schedules; the last value repeats if you run out.
  • Caption - max new tokens - the shared token budget all witnesses get.
  • Distiller - model / Validator - model - the Ollama tags for Pass B and Pass C (dropdowns from config/captionforge_ollama_models.json, with Custom for arbitrary tags).
  • Final - caption style - narrative, comma, or both.
  • LoRA - trigger word / LoRA - user caption anchor - threaded through distiller and validator.

Outputs: single_image (your IMAGE passthrough), pipeline_plan, and pipeline_plan_json.

Install

Same pack install as everything else:

git clone https://github.com/Damkohler/CaptionForge.git ComfyUI/custom_nodes/CaptionForge

Restart ComfyUI. The planner itself loads no models, but the run it plans needs Ollama up with the distiller/validator tags pulled (ollama pull mistral-small:24b, ollama pull gemma4:26b).

Common issues

This node is where runs get misconfigured, so read your settings like a checklist: witness counts, seed mode, output folder, trigger word. If a caption node seems to ignore your standalone settings, check that the planner's overrides aren't stomping them - that's by design. And if a run produces nothing, the usual cause is a typo in Input - image path or a glob that matches zero files. Start with one image and low witness counts before you point it at a 10,000-image dataset, because the pack is unapologetically heavy.

CategoryCaptioning/CaptionForge

Inputs (45)

NameTypeDefaultDescription
Planner - enabledBOOLEANtrueEnable CaptionForge Pipeline Planner mode. If disabled, this node passes through the optional IMAGE and emits an empty/falsy plan so downstream nodes can run in standalone mode without GUI bypassing.
Input - image pathSTRINGImage file or image folder/root for ordinary folder/file workflows. This is also used as the validator image root. For quick single-image workflows, connect IMAGE to the optional Input - single image socket.
Input - recursiveBOOLEANtrueWhether captioning nodes should recurse when Input - image path is a folder.
Input - filename globSTRING*Filename glob for folder captioning, e.g. *.png, *.jpg, or *.
Output - folderSTRING/tmp/ComfyUI/output/CaptionForgeOutput root folder. CaptionForge creates a run-specific working directory inside this folder for JSON/JSONL/audit artifacts. Final TXT sidecars are written beside their resolved source images.
Output - run nameSTRINGcaptionforge_runRun-root used to name config JSON, output path JSON, caption JSONL, distiller JSONL, prompt JSONLs, validator JSONL, and final JSONL.
Output - overwrite outputsBOOLEANtrueOverwrite generated run artifacts during capstone execution.
LoRA - trigger wordSTRINGOptional shared LoRA trigger token/string preserved through the pipeline.
LoRA - user caption anchorSTRINGOptional user style/identity anchor passed to distiller and validator.
Caption - Joy runs/imageCOMBO2Joy Caption runs per image. Set to Disabled to omit Joy from this run. Dropdown is capped at 5 to prevent accidental giant runs.
Caption - Qwen runs/imageCOMBO2Qwen Caption runs per image. Set to Disabled to omit Qwen from this run. Dropdown is capped at 5 to prevent accidental giant runs.
Caption - Ollama runs/imageCOMBODisabledOllama Caption runs per image for each connected JLC CaptionForge Ollama Caption node. The actual Ollama model tag is selected in each Ollama Caption node. Connecting multiple Ollama Caption nodes multiplies runtime, memory pressure, and raw-caption count.
Caption - base seedINT-1-1–4294967295Base seed for caption generation. -1 means unseeded when supported.
Caption - seed modeCOMBOfixed4 options: fixed, increment, decrement, random
Caption - temperature scheduleSTRING0.90Comma-separated caption temperature schedule; final value repeats if needed.
Caption - top p scheduleSTRING0.60
Caption - top k scheduleSTRING80
Caption - max image sizeINT10240–4096
Caption - max new tokensINT600016–12000
Distiller - modelCOMBOmistral-small:24bConcrete Ollama text model tag for the distiller. Use custom to enter any other installed Ollama text model.
Distiller - custom Ollama modelSTRINGUsed only when Distiller - model is custom, e.g. my-model:latest.
Distiller - base seedINT-1-1–4294967295Base seed for the distiller. -1 means omit seed.
Distiller - seed modeCOMBOfixed4 options: fixed, increment, decrement, random
Distiller - strategyCOMBOsingle_pass2 options: single_pass, by_model_then_global
Distiller - max caption chars for LLMINT15360–12000
Distiller - num predictINT309664–12000
Distiller - temperatureFLOAT0.240–2
Distiller - top pFLOAT0.900–1
Distiller - top kINT600–500
Distiller - write prompt JSONLBOOLEANfalse
Distiller - preserve raw responseBOOLEANfalse
Validator - modelCOMBOgemma4:26bConcrete Ollama vision model tag for the image-aware validator. Use custom to enter any other installed Ollama vision model.
Validator - custom Ollama modelSTRINGUsed only when Validator - model is custom, e.g. gemma4:e4b or another installed VLM tag.
Validator - base seedINT-1-1–4294967295Base seed for the VLM validator. -1 means omit seed.
Validator - seed modeCOMBOfixed4 options: fixed, increment, decrement, random
Validator - num predictINT220064–12000
Validator - temperatureFLOAT0.000–2
Validator - top pFLOAT0.920–1
Validator - top kINT800–500
Validator - write prompt JSONLBOOLEANfalse
Validator - preserve raw VLM responseBOOLEANfalse
Final - caption styleCOMBOnarrative3 options: narrative, comma, both
Final - write TXT sidecarsBOOLEANtrueFinal TXT sidecars use the associated image filename stem and .txt extension.
Final - write JSONLBOOLEANtrue
Input - single imageoptIMAGEOptional IMAGE passthrough for quick single-image workflows. The planner does not process this image; it simply returns it as output.

Outputs (3)

NameTypeDescription
single_imageIMAGE
pipeline_planCAPTIONFORGE_PIPELINE_PLAN
pipeline_plan_jsonSTRING