DIGIT Batch Caption
Caption a whole image folder for LoRA training without touching a file
- log
- captioned_count
- folder_path
Captioning is the step everyone keeps meaning to do and never actually does. DIGIT Batch Caption is the node that fixes that: point it at a folder of images, and it sends each one to a Gemini model through Google Vertex AI, then writes a matching .txt sidecar next to every image. No GPU, no Python script, no clipboard gymnastics. Set it going, come back to a dataset that's actually ready to train on.
You'll reach for this when you're assembling a LoRA dataset. Caption quality genuinely beats most training knobs - describe what you want to stay variable, leave the fixed stuff (identity, style) out of the captions. What this node writes are natural-language captions, which is exactly what the LLM-encoder bases (Flux, Qwen-Image, Z-Image, Krea 2) want. If you're training an Illustrious or Pony character, a Danbooru-tags style (booru_tags) is in the dropdown too, so it covers both worlds.
How it works
The node scans image_folder, resizes each image down to max_dimension (default 2048) before sending it, base64-encodes it, and fires it at the Gemini model you picked. The caption comes back and lands in a .txt beside the image. Images that already have a caption are skipped unless overwrite is on - which makes the node resumable, a genuinely nice touch when a 500-image folder hiccups halfway.
The few inputs you actually set:
image_folder- where your images live. Required.caption_style- defaults totraining_detailed, which is the right starting point for a training set.training_concisefor shorter,booru_tagsfor the tagged-anime lineage.caption_length-short,medium,long, orany. Defaultlong.overwrite- off by default, so re-runs skip what's done. Flip it on only if you changed the prompt and want everything redone.trigger_word- connect this from a DIGIT LoRA Loader output (or type it) and every caption gets the trigger prepended. One less thing to forget.prefix_text/suffix_text- consistent style or quality tags injected at the start or end of every caption.
Also worth knowing: max_tokens (1024 default), temperature (0.4 default, low is good for captions), model (gemini-2.5-flash default, up to gemini-3.1-pro-preview if you want the pricier, more detailed pass), and delay_seconds (0.5) which paces API calls so you don't hammer rate limits.
Outputs are log (per-file status), captioned_count, and folder_path - wire that last one into DIGIT Caption Viewer to actually eyeball what Gemini wrote before you train on it. That QA loop is worth the ten seconds.
Installing it
This node ships in the digit-comfyui pack. ComfyUI Manager → search comfyui-digit → Install → restart. Manual route:
cd ComfyUI/custom_nodes
git clone https://github.com/thedepartmentofexternalservices/comfyui-digit.git
cd comfyui-digit
pip install -r requirements.txt
The Google dependency is the real install: gcloud plus a Vertex AI–enabled project, authenticated with gcloud auth application-default login. Do that once and every Gemini node in the pack picks up the credentials automatically. No API key to manage - that's the pack's whole pitch.
Where people get burned
The usual failure is auth: the node auto-detects the project from gcp_project_id, the DIGIT_GCP_PROJECT env var, or GCP metadata, and if none of those are set you get a credentials error. Sort the GCP setup first and everything else is boring. And remember what you're signing up for - this is a hosted captioner, so your images leave the machine and Google's per-call pricing applies. A big folder is real money, and GCP's terms allow prompt logging when abuse filters trip. For a few hundred images it's the best time-per-dollar captioning around; for a 10,000-image style run, budget first.
Inputs (16)
| Name | Type | Default | Description |
|---|---|---|---|
| image_folder | STRING | Path to folder containing images to caption. | |
| model | COMBO | gemini-2.5-flash | 9 options: gemini-2.5-pro, gemini-2.5-flash, gemini-2.5-flash-lite, gemini-2.0-flash, gemini-2.0-flash-lite, gemini-3.1-pro-preview, +3 |
| caption_style | COMBO | training_detailed | 7 options: descriptive_formal, descriptive_casual, training_detailed, training_concise, booru_tags, prompt_style, +1 |
| caption_length | COMBO | long | 4 options: short, medium, long, any |
| overwrite | BOOLEAN | false | Overwrite existing .txt caption files. If false, skips images that already have captions. |
| gcp_project_id | STRING | GCP project ID. Auto-detected from DIGIT_GCP_PROJECT env var or GCP metadata. | |
| gcp_region | STRING | GCP region. Auto-detected from DIGIT_GCP_REGION env var or GCP metadata. Defaults to 'global'. | |
| trigger_wordopt | STRING | Trigger word prepended to each caption. Connect from LoRA Loader or type manually. Leave disconnected to skip. | |
| prefix_textopt | STRING | Text injected at the START of every caption (after trigger word). Useful for consistent style tags. | |
| suffix_textopt | STRING | Text injected at the END of every caption. Useful for quality tags or consistent endings. | |
| custom_promptopt | STRING | Custom prompt used when caption_style is 'custom'. Also appended as extra instructions for other styles. | |
| system_promptopt | STRING | Optional system prompt override. If empty, a default captioning system prompt is used. | |
| max_tokensopt | INT | 102464–8192 | — |
| temperatureopt | FLOAT | 0.400–2 | — |
| max_dimensionopt | INT | 2048512–4096 | Resize images to this max dimension before sending to API (saves bandwidth). |
| delay_secondsopt | FLOAT | 0.50–10 | Delay between API calls to avoid rate limiting. |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| log | STRING | — |
| captioned_count | INT | — |
| folder_path | STRING | — |