Nodes/DIGIT Nodes/DIGIT Captioner
ComfyUI Node

DIGIT Captioner

The Gemini captioner that lives inside the LoRA training pipeline

By thedepartmentofexternalservices·Created 7 months ago·Updated 2 months ago· 0
DIGIT Captioner
    • report
    • last_caption
    • captioned_count
    ◄dataset_path►
    ◄actioncaption_uncaptioned►
    ◄caption_preset►
    ◄system_promptYou are an expert image captioner for AI training datasets. Describe the image in detail, focusing on subject, composition, lighting, colors, style, and mood. Be specific and descriptive. Output only the caption, no preamble.►
    ◄prompt_templateDescribe this image in detail for AI training:►
    ◄modelgemini-2.5-flash►
    ◄temperature0.40►
    ◄max_tokens300►
    ◄overwritefalse►
    ◄caption_ext.txt►
    ◄single_image_path►
    ◄gcp_project_id►
    ◄gcp_region►

    Where DIGIT Batch Caption is a standalone "caption this folder" tool, the DIGIT Captioner is the version that lives inside the pack's LoRA training suite. Same underlying idea - Gemini writes captions for your dataset through Vertex AI - but organized around training runs: it only captions what's missing, can re-caption everything, handles a single image, and reads its captioning recipe from a saved preset.

    You'll reach for this when your workflow is already built on the DIGIT training nodes and you want captioning to behave like a step in the pipeline rather than a one-off job. The default action is caption_uncaptioned, which is the right instinct: caption only the images that don't have a sidecar yet, and never touch the ones you already captioned by hand. Re-running is cheap because it does nothing.

    How it works

    It scans dataset_path for images, checks for existing caption files (caption_ext, default .txt), and sends each image to Gemini with your system prompt and prompt template. The actions tell it which subset to work on:

    • caption_uncaptioned (default) - only images missing a caption.
    • caption_all - every image, respecting overwrite.
    • caption_single - one image via single_image_path, for spot-fixing.
    • recaption_all - force everything, ignoring existing captions.
    • preview - show what it would write without committing.

    The inputs that shape output: system_prompt (a solid detailed-captioning default ships with it), prompt_template, model (gemini-2.5-flash default), temperature (0.4 - keep it low, captioning wants consistency not creativity), max_tokens (300), and overwrite. The caption_preset field is where you load a saved recipe from the DIGIT Caption Preset Manager, which is the neat part: tune a preset once, reuse it across datasets instead of re-typing system prompts.

    Outputs: report (a summary of what happened), last_caption (the most recent caption written, handy for spot-checks), and captioned_count.

    Installing it

    Part of digit-comfyui:

    cd ComfyUI/custom_nodes
    git clone https://github.com/thedepartmentofexternalservices/comfyui-digit.git
    cd comfyui-digit
    pip install -r requirements.txt
    

    Or ComfyUI Manager → search comfyui-digit → install → restart. The full training suite wants the extra deps too (pip install -r requirements-training.txt) if you're going all the way to the trainer, but the Captioner itself only needs the base requirements plus your GCP auth: gcloud auth application-default login and a Vertex AI–enabled project, auto-detected or set via gcp_project_id.

    Notes from the trenches

    Two things worth knowing from how people actually caption. First, natural-language captions from Gemini are the right format for the modern LLM-encoder bases (Flux, Qwen-Image, Z-Image) - write what should vary, leave identity features undescribed, and this node's detailed default prompt is already pointed that way. Second, don't blow up temperature; a 0.4 default exists because a captioner that gets creative is a captioner that invents details your model will bake in. And since this is a hosted captioner, budget for per-call GCP billing on big datasets - caption_uncaptioned is also your friend there, because re-runs genuinely cost nothing.

    CategoryDIGIT

    Inputs (13)

    NameTypeDefaultDescription
    dataset_pathSTRING—
    actionCOMBOcaption_uncaptioned5 options: caption_all, caption_uncaptioned, caption_single, recaption_all, preview
    caption_presetoptSTRING—
    system_promptoptSTRINGYou are an expert image captioner for AI training datasets. Describe the image in detail, focusing on subject, composition, lighting, colors, style, and mood. Be specific and descriptive. Output only the caption, no preamble.—
    prompt_templateoptSTRINGDescribe this image in detail for AI training:—
    modeloptCOMBOgemini-2.5-flash4 options: gemini-2.5-flash, gemini-2.5-pro, gemini-2.0-flash, gemini-2.0-flash-lite
    temperatureoptFLOAT0.400–2—
    max_tokensoptINT30050–2000—
    overwriteoptBOOLEANfalse—
    caption_extoptSTRING.txt—
    single_image_pathoptSTRING—
    gcp_project_idoptSTRING—
    gcp_regionoptSTRING—

    Outputs (3)

    NameTypeDescription
    reportSTRING—
    last_captionSTRING—
    captioned_countINT—