JLC CaptionForge Qwen Caption
Qwen captioning from Hugging Face
- image
- pipeline_plan
- template_options
- image_out
- pipeline_plan_out
- template_options_out
- caption
- resolved_prompt
Joy is the crowd favorite, but it's not the only game in town - and CaptionForge's whole premise is that a second, independent captioning voice catches what the first one misses. JLC CaptionForge Qwen Caption is that second voice: a Qwen-family vision-language model loaded from Hugging Face, run as a Pass A witness alongside Joy (and optionally Ollama). It's structurally a sibling of the Joy node - same template path, same outputs, same standalone-or-pipeline split - but with a different model and a few different defaults.
The practical argument for it: Joy and Qwen have different blind spots, and when a distiller later merges their accounts, details that appear in both get reinforced while contradictions get flagged. The README is honest that model choices matter a lot and that Qwen's value is mostly as a complementary voice. If you only run one captioner, Joy is the community default; if you run two, this is the natural second.
How it works
It's a Python/Hugging Face engine, so weights load into your ComfyUI process (managed through the CaptionForge model cache, with eviction before Ollama stages run). Models are expected under ComfyUI/models/LLM/JLC_QwenCaption/ and auto-download on first use unless download_probe_only is on. The model dropdown ships five variants, from the lightweight Qwen2.5-VL-3B-Instruct up to Qwen2.5-VL-7B-Instruct, plus community finetunes like the Unredacted and abliterated NSFW-caption variants - which tells you what kind of datasets this pack expects.
qwen_quantization defaults to Balanced (8-bit), bitsandbytes load-time quantization that's genuinely worth keeping for the 7B variants on 16 GB cards. The sampling defaults differ from Joy's too - repetition_penalty sits at 1.08 here versus Joy's 1.0, a small nudge against looped captions.
Inputs and outputs that matter
- model - pick your Qwen variant; bigger isn't automatically better on a 16 GB card.
- qwen_quantization -
Balanced (8-bit)vsDefault. - caption_type / caption_length - the template path (Descriptive, LoRA Literal, Taggy, Style Focus, SFW Character Caption…).
- system_prompt / custom_prompt - same structure as Joy;
custom_prompt_modeoverrides the template path when enabled. - max_new_tokens / temperature / top_p / top_k - standalone sampling, overridden by the Pipeline Planner when a plan is connected.
- forbidden_phrases / replace_pairs - the shared cleanup filter.
Outputs are the familiar five: caption, resolved_prompt, image_out, pipeline_plan_out, template_options_out. caption is what feeds your dataset or the capstone's JSONL; resolved_prompt shows you the exact prompt that went to the model.
Install
Pack install once:
git clone https://github.com/Damkohler/CaptionForge.git ComfyUI/custom_nodes/CaptionForge
Restart, then pip install -e . in the folder if your ComfyUI env lacks the deps (torch, transformers, accelerate, huggingface-hub, pillow, numpy, safetensors, qwen-vl-utils). pip install bitsandbytes enables the 8-bit path. The model itself auto-downloads into ComfyUI/models/LLM/JLC_QwenCaption/ - no manual weight step.
Common issues
First run is a multi-GB Hugging Face download, so either be patient or dry-run with download_probe_only. If you OOM on a 7B model, check that qwen_quantization is actually set to Balanced (8-bit) - that's the single biggest lever. And expect VRAM handoffs in full-pipeline runs: Qwen and Joy share the process-local model cache and get evicted before Ollama stages take over, which is normal behavior, not a leak.
Inputs (23)
| Name | Type | Default | Description |
|---|---|---|---|
| model | COMBO | Qwen2.5-VL-7B-Instruct | Qwen vision-language model. Models are loaded from ComfyUI/models/LLM/JLC_QwenCaption/. Missing models may be downloaded automatically unless download_probe_only is enabled. |
| qwen_quantization | COMBO | Balanced (8-bit) | Qwen model load mode. Balanced (8-bit) uses bitsandbytes 8-bit loading to reduce VRAM pressure, especially for Qwen2.5-VL 7B variants. |
| keep_loaded | BOOLEAN | true | Keep the model cached after captioning for faster repeated runs. CaptionForge cache policy may still evict it when another caption model must load. |
| caption_template_mode | BOOLEAN | true | Use the structured CaptionForge template path: caption_type, caption_length, and optional Template Options from the template_options pin. If custom_prompt_mode is also enabled, custom_prompt_mode takes precedence. |
| caption_type | COMBO | LoRA Literal | Caption template style used when caption_template_mode is active. |
| caption_length | COMBO | any | Target caption length used when caption_template_mode is active. |
| custom_prompt_mode | BOOLEAN | false | Use custom_prompt when non-empty, otherwise use prompt_preset. This overrides caption_template_mode when both toggles are enabled. |
| prompt_preset | COMBO | default_literal | Built-in prompt preset used only in custom_prompt_mode when custom_prompt is blank. |
| system_prompt | STRING | You are a helpful image-captioning assistant. Describe only what is visible in the image. Do not invent unseen context. | Qwen engine accepts a single prompt string, so this system instruction is folded above the resolved caption prompt. Kept next to custom_prompt for clarity. |
| custom_prompt | STRING | Custom prompt used only when custom_prompt_mode is enabled. Overrides prompt_preset when non-empty. | |
| max_new_tokens | INT | 38416–4096 | Standalone token budget. When a Pipeline Planner is connected, this is overridden by the Planner's shared max_new_tokens. |
| temperature | FLOAT | 0.750–2 | Standalone sampling temperature. When a Pipeline Planner is connected, this is overridden by the Planner temperature schedule. |
| top_p | FLOAT | 0.900–1 | Standalone top-p sampling value. When a Pipeline Planner is connected, this is overridden by the Planner top-p schedule. |
| top_k | INT | 500–500 | Standalone top-k sampling limit. When a Pipeline Planner is connected, this is overridden by the Planner top-k schedule. |
| repetition_penalty | FLOAT | 1.081–2 | Penalty applied to repeated tokens. Kept with the core captioning parameters. This is not currently overridden by the Pipeline Planner. |
| max_size | INT | 10240–4096 | Maximum longest-side image size for standalone captioning. The image is resized in memory only. Pipeline Planner overrides this in planned runs. |
| forbidden_phrases | STRING | Optional cleanup filter: remove lines/captions containing any listed phrase, one per line. | |
| replace_pairs | STRING | Optional cleanup replacements, one per line: old=>new. | |
| download_probe_only | BOOLEAN | false | At the very bottom by design. Probe/download lightweight model metadata only, then return a status message without captioning. |
| imageopt | IMAGE | Image or batch of images to caption. The image is passed through unchanged for clean node-to-node pipeline chaining. | |
| pipeline_planopt | CAPTIONFORGE_PIPELINE_PLAN | Connect the CaptionForge Pipeline Planner output here. When connected, this node switches into Pass A evidence mode: Planner image routing, per-run seeds, sampling schedules, shared output paths, and internal JSONL evidence append. | |
| template_optionsopt | CAPTIONFORGE_EXTRA_OPTIONS | Connect the CaptionForge Template Options node here. Works in standalone and Pipeline modes. This is the only source for template modifiers and name input. | |
| seedopt | INT | Optional standalone seed input. Ignored when a Pipeline Planner supplies a seed schedule. |
Outputs (5)
| Name | Type | Description |
|---|---|---|
| image_out | IMAGE | — |
| pipeline_plan_out | CAPTIONFORGE_PIPELINE_PLAN | — |
| template_options_out | CAPTIONFORGE_EXTRA_OPTIONS | — |
| caption | STRING | — |
| resolved_prompt | STRING | — |