ComfyUI Node
JLC CaptionForge Node
A ComfyUI node in Captioning/CaptionForge with 49 inputs and 5 outputs.
JLC CaptionForge Node
- Input - single image
- pipeline_plan
- natural_captions
- taggy_captions
- final_jsonl_records
- output_paths_json
- status
◄Input - captions JSONL►
◄Input - image path►
◄Input - include caption familiesjoy,qwen,ollama►
◄Input - max captions per family5►
◄Input - max total captions20►
◄Output - folder►
◄Output - run namecaptionforge_run►
◄Output - overwrite outputstrue►
◄Ollama - URLhttp://127.0.0.1:11434►
◄Ollama - keep loadedtrue►
◄Ollama - request timeout seconds1800►
◄LoRA - trigger word►
◄LoRA - user caption anchor►
◄Fat Draft - modelmistral-small:24b►
◄Fat Draft - custom Ollama model►
◄Fat Draft - prompt/no_think
You are a detail-preserving caption merger for LoRA dataset preparation.
You receive multiple captions of the same image. You do NOT see the image.
Task:
Merge all non-contradictory caption details into one deliberately over-complete draft caption.
Rules:
- Do not validate against the image.
- Do not decide that details are false just because they appear once.
- Do not summarize aggressively.
- Preserve concrete details from all captions.
- Split contradictions by choosing cautious wording or listing the alternative only when needed.
- Prefer specific visual language over generic language.
- Keep visible body, clothing, material, accessory, color, pose, lighting, style, and framing details.
- Preserve doll-like, glossy/plastic-like, material, garment-construction, body-shape, and facial-feature details when present.
- Use neutral dataset-caption language, including visible sensual styling or revealing clothing when present.
- Do not add details absent from the captions.
- Treat subject names or trigger-like identity tokens as optional identity labels. Preserve them only when they appear consistently in the captions; do not let them replace visible description.
- Output only one paragraph, no notes, no JSON.►
◄Fat Draft - base seed1►
◄Fat Draft - seed modefixed►
◄Fat Draft - max caption chars1536►
◄Fat Draft - max new tokens5000►
◄Fat Draft - temperature0.12►
◄Fat Draft - top p0.88►
◄Fat Draft - top k50►
◄Validator - modelgemma4:26b►
◄Validator - custom Ollama model►
◄Validator - system prompt/no_think
You are a direct image validation engine. Inspect the image and answer only with the requested caption.►
◄Validator - prompt/no_think
Look at the image and validate this draft caption.
Task:
Return a corrected caption paragraph that keeps only image-supported details.
Rules:
- Output only the corrected caption.
- One paragraph.
- No reasoning, no notes, no JSON.
- Keep all true visible details from the draft.
- Delete unsupported details.
- Correct small visible errors.
- Do not add new details unless needed to correct an error already present.
- Preserve useful LoRA details: subject, face, hair, eyes, makeup, lips, skin texture, pose, body shape, outfit, accessories, materials, colors, lighting, background, framing, and visual style.
- Visible sensual styling, revealing clothing, cleavage, thighs, bare skin, swimwear, lingerie, or body-shape details may be described neutrally when present.
- Do not invent hidden anatomy, unseen clothing, explicit acts, or details contradicted by the image.►
◄Validator - base seed1►
◄Validator - seed modefixed►
◄Validator - max new tokens5000►
◄Validator - temperature0.05►
◄Validator - top p0.88►
◄Validator - top k50►
◄Formatter - modelmistral-small:24b►
◄Formatter - custom Ollama model►
◄Formatter - prompt/no_think
You are a LoRA caption format converter.
The validated paragraph is already the natural-language final caption. Do not rewrite it.
Task:
Create one TAGGY caption from the validated paragraph.
Rules:
- Output only the taggy comma-separated caption.
- Use only details already present in the validated paragraph.
- Preserve concrete LoRA-useful details.
- Do not add new details.
- Do not mention this process.
- Do not output markdown.
- Do not include a TAGGY: label.
- Keep the result as a comma-separated list, not full prose.►
◄Formatter - base seed1►
◄Formatter - seed modefixed►
◄Formatter - max new tokens3200►
◄Formatter - temperature0.12►
◄Formatter - top p0.88►
◄Formatter - top k50►
◄Audit - write prompt JSONLfalse►
◄Audit - preserve raw responsesfalse►
◄Final - TXT export formatnatural►
◄Final - write TXT sidecarstrue►
◄Final - write JSONLtrue►
CategoryCaptioning/CaptionForge
Inputs (49)
| Name | Type | Default | Description |
|---|---|---|---|
| Input - captions JSONL | STRING | Pass A raw caption JSONL produced by CaptionForge Caption nodes. | |
| Input - image path | STRING | Image file/folder root used by the VLM validator to resolve source images. | |
| Input - include caption families | STRING | joy,qwen,ollama | Comma-separated model_family values to use from Pass A. Use all or * to include everything. |
| Input - max captions per family | INT | 50–50 | Maximum selected Pass A captions per model family. 0 means no per-family cap. |
| Input - max total captions | INT | 200–100 | Maximum selected Pass A captions per image. 0 means no total cap. |
| Output - folder | STRING | Output folder. Planner value overrides this when pipeline_plan is connected. | |
| Output - run name | STRING | captionforge_run | Run-root used for B/C/D/E JSONL and TXT artifacts. |
| Output - overwrite outputs | BOOLEAN | true | — |
| Ollama - URL | STRING | http://127.0.0.1:11434 | Local Ollama server URL. |
| Ollama - keep loaded | BOOLEAN | true | Pass keep_alive to Ollama. Ollama ultimately owns model residency. |
| Ollama - request timeout seconds | INT | 180010–7200 | HTTP patience for Ollama calls. This does not affect caption quality. |
| LoRA - trigger word | STRING | — | |
| LoRA - user caption anchor | STRING | — | |
| Fat Draft - model | COMBO | mistral-small:24b | 6 options: mistral-small:24b, VladimirGav/gemma4-26b-16GB-VRAM-Uncensored, deepseek-r1:32b, tarruda/neuraldaredevil-8b-abliterated:fp16, gpt-oss:20b, Custom |
| Fat Draft - custom Ollama model | STRING | Used only when Fat Draft - model is Custom. | |
| Fat Draft - prompt | STRING | /no_think You are a detail-preserving caption merger for LoRA dataset preparation. You receive multiple captions of the same image. You do NOT see the image. Task: Merge all non-contradictory caption details into one deliberately over-complete draft caption. Rules: - Do not validate against the image. - Do not decide that details are false just because they appear once. - Do not summarize aggressively. - Preserve concrete details from all captions. - Split contradictions by choosing cautious wording or listing the alternative only when needed. - Prefer specific visual language over generic language. - Keep visible body, clothing, material, accessory, color, pose, lighting, style, and framing details. - Preserve doll-like, glossy/plastic-like, material, garment-construction, body-shape, and facial-feature details when present. - Use neutral dataset-caption language, including visible sensual styling or revealing clothing when present. - Do not add details absent from the captions. - Treat subject names or trigger-like identity tokens as optional identity labels. Preserve them only when they appear consistently in the captions; do not let them replace visible description. - Output only one paragraph, no notes, no JSON. | Instructions for the text-only fat draft LLM. Captions are appended automatically. |
| Fat Draft - base seed | INT | 1-1–4294967295 | — |
| Fat Draft - seed mode | COMBO | fixed | 4 options: fixed, increment, decrement, random |
| Fat Draft - max caption chars | INT | 15360–12000 | — |
| Fat Draft - max new tokens | INT | 500064–12000 | Maps to Ollama num_predict. |
| Fat Draft - temperature | FLOAT | 0.120–2 | — |
| Fat Draft - top p | FLOAT | 0.880–1 | — |
| Fat Draft - top k | INT | 500–500 | — |
| Validator - model | COMBO | gemma4:26b | 4 options: gemma4:26b, qwen3.6:35B-A3B, huihui_ai/gemma-4-abliterated:26b, Custom |
| Validator - custom Ollama model | STRING | Used only when Validator - model is Custom. | |
| Validator - system prompt | STRING | /no_think You are a direct image validation engine. Inspect the image and answer only with the requested caption. | System prompt for the image-aware VLM validator. |
| Validator - prompt | STRING | /no_think Look at the image and validate this draft caption. Task: Return a corrected caption paragraph that keeps only image-supported details. Rules: - Output only the corrected caption. - One paragraph. - No reasoning, no notes, no JSON. - Keep all true visible details from the draft. - Delete unsupported details. - Correct small visible errors. - Do not add new details unless needed to correct an error already present. - Preserve useful LoRA details: subject, face, hair, eyes, makeup, lips, skin texture, pose, body shape, outfit, accessories, materials, colors, lighting, background, framing, and visual style. - Visible sensual styling, revealing clothing, cleavage, thighs, bare skin, swimwear, lingerie, or body-shape details may be described neutrally when present. - Do not invent hidden anatomy, unseen clothing, explicit acts, or details contradicted by the image. | Instructions for the image-aware VLM validator. The fat draft is appended automatically. |
| Validator - base seed | INT | 1-1–4294967295 | — |
| Validator - seed mode | COMBO | fixed | 4 options: fixed, increment, decrement, random |
| Validator - max new tokens | INT | 500064–12000 | Maps to Ollama num_predict. |
| Validator - temperature | FLOAT | 0.050–2 | — |
| Validator - top p | FLOAT | 0.880–1 | — |
| Validator - top k | INT | 500–500 | — |
| Formatter - model | COMBO | mistral-small:24b | 5 options: mistral-small:24b, VladimirGav/gemma4-26b-16GB-VRAM-Uncensored, gpt-oss:20b, deepseek-r1:32b, Custom |
| Formatter - custom Ollama model | STRING | Used only when Formatter - model is Custom. | |
| Formatter - prompt | STRING | /no_think You are a LoRA caption format converter. The validated paragraph is already the natural-language final caption. Do not rewrite it. Task: Create one TAGGY caption from the validated paragraph. Rules: - Output only the taggy comma-separated caption. - Use only details already present in the validated paragraph. - Preserve concrete LoRA-useful details. - Do not add new details. - Do not mention this process. - Do not output markdown. - Do not include a TAGGY: label. - Keep the result as a comma-separated list, not full prose. | Instructions for the text-only taggy formatter. The validated paragraph is appended automatically. |
| Formatter - base seed | INT | 1-1–4294967295 | — |
| Formatter - seed mode | COMBO | fixed | 4 options: fixed, increment, decrement, random |
| Formatter - max new tokens | INT | 320064–12000 | Maps to Ollama num_predict. |
| Formatter - temperature | FLOAT | 0.120–2 | — |
| Formatter - top p | FLOAT | 0.880–1 | — |
| Formatter - top k | INT | 500–500 | — |
| Audit - write prompt JSONL | BOOLEAN | false | — |
| Audit - preserve raw responses | BOOLEAN | false | — |
| Final - TXT export format | COMBO | natural | 3 options: natural, taggy, both_separate |
| Final - write TXT sidecars | BOOLEAN | true | — |
| Final - write JSONL | BOOLEAN | true | — |
| Input - single imageopt | IMAGE | Optional IMAGE passthrough/reference for planned single-image workflows. | |
| pipeline_planopt | CAPTIONFORGE_PIPELINE_PLAN | Connect the CaptionForge Pipeline Planner pipeline_plan output here. |
Outputs (5)
| Name | Type | Description |
|---|---|---|
| natural_captions | STRING | — |
| taggy_captions | STRING | — |
| final_jsonl_records | STRING | — |
| output_paths_json | STRING | — |
| status | STRING | — |