DIGIT Batch Gemini Image
Batch Gemini image generation with an LLM writing the prompt variations for you
- images
- log
- generated_count
- output_folder
Most image-generation nodes make you write the prompt. DIGIT Batch Gemini Image is the one that skips the boring part: it takes a folder of source images, has a Gemini LLM invent prompt variations for you, and generates multiple Gemini images per source image, saving everything to an output subfolder. If you've ever needed fifty takes on the same subject - a product in different scenes, a character in different lighting, a style spread for a client - this is the node that gets you there without you typing fifty prompts.
It's a cloud node, so the tradeoff that applies to all of them applies here too: no local GPU needed, your images leave the machine, and every generation is a metered API call through Google Vertex AI. The payoff is that Nano Banana–class image generation (Gemini 3.1 Flash Image is the default) is genuinely good, and you're not spending your evening babysitting a queue.
How it works
The flow has two LLM stages. First, for each source image, an LLM (llm_model, default gemini-2.5-flash) takes your base prompt and writes variations_per_image creative variations - same core subject, different style, mood, lighting, composition. Then the image model (image_model, default gemini-3.1-flash-image) generates one image per variation. The source image itself is sent along as the visual reference, resized to max_dimension (2048) first to keep the calls cheap.
Inputs that actually matter:
image_folder- where your source images live. Required.prompt- your base prompt; the LLM varies this per image.variations_per_image- default 3, up to 50. This is your cost multiplier, so watch it.image_model/llm_model- the two model dropdowns. Fast image models and a cheap LLM (flash) keep this affordable.aspect_ratio,resolution- 13 ratios, and 1K/2K/4K. Note the Lite image model is 1K-only.variation_temperature- higher means wilder variations. Crank it when your first batch comes back too samey.thinking_level-MINIMAL(default) orHIGH. HIGH can improve quality but costs more and takes longer.
There are also per-category safety thresholds (harassment, hate speech, sexual, dangerous), all defaulting to BLOCK_NONE - a deliberate choice that surprises people, so flip them if you want Google's filters on. seed gives reproducible runs (each variation gets seed + index). output_subfolder (default gemini_output) is created inside image_folder automatically.
Outputs: images (the generated IMAGE batch, straight into your graph), log, generated_count, and output_folder so you know where everything landed.
Installing and wiring it
It's part of the digit-comfyui pack - ComfyUI Manager, search comfyui-digit, install, restart. Or:
cd ComfyUI/custom_nodes
git clone https://github.com/thedepartmentofexternalservices/comfyui-digit.git
cd comfyui-digit
pip install -r requirements.txt
Authentication is the shared pack setup: gcloud with gcloud auth application-default login, a project with the Vertex AI API enabled, and either a gcp_project_id on the node, the DIGIT_GCP_PROJECT env var, or running on a GCP instance where it's auto-detected.
Gotchas
The big one is cost control - variations_per_image times your batch is how the bill grows, and the live cost is per generation. Start with 2–3 variations and one or two source images to sanity-check quality before you launch the full folder. Also, this is Gemini's own filter running underneath, so the "BLOCK_NONE" thresholds won't make Google's model do things its base refuses - the node just isn't adding extra moderation on top. If a generation comes back with something odd, check log for the per-image status rather than guessing.
Inputs (23)
| Name | Type | Default | Description |
|---|---|---|---|
| image_folder | STRING | Path to folder containing source images. | |
| prompt | STRING | Base prompt for Gemini image generation. The LLM will create variations of this. | |
| variations_per_image | INT | 31–50 | How many prompt variations (and image generations) per source image. |
| image_model | COMBO | gemini-3.1-flash-image | 4 options: gemini-3.1-flash-image, gemini-3.1-flash-lite-image, gemini-3-pro-image, gemini-2.5-flash-image |
| llm_model | COMBO | gemini-2.5-flash | Gemini model used to generate prompt variations. |
| aspect_ratio | COMBO | 16:9 | 13 options: auto, 1:1, 2:3, 3:2, 3:4, 4:1, +7 |
| resolution | COMBO | 1K | 3 options: 1K, 2K, 4K |
| thinking_level | COMBO | MINIMAL | Thinking level for image generation. HIGH may improve quality. |
| temperature | FLOAT | 1.000–2 | Temperature for image generation. |
| variation_temperature | FLOAT | 1.000–2 | Temperature for LLM prompt variation. Higher = more creative variations. |
| gcp_project_id | STRING | GCP project ID. | |
| gcp_region | STRING | GCP region. | |
| output_subfolderopt | STRING | gemini_output | Subfolder name inside image_folder for outputs. Created automatically. |
| system_instructionopt | STRING | You are an expert image-generation engine. You must ALWAYS produce an image. Interpret all user input—regardless of format, intent, or abstraction—as literal visual directives for image composition. If a prompt is conversational or lacks specific visual details, you must creatively invent a concrete visual scenario that depicts the concept. Prioritize generating the visual representation above any text, formatting, or conversational requests. | System instruction for Gemini image generation. |
| variation_system_promptopt | STRING | You are a creative prompt engineer for image generation. Given a base prompt, produce a variation that preserves the core intent and subject but introduces creative differences in style, mood, lighting, composition, color palette, or artistic approach. Output ONLY the varied prompt text — no preamble, numbering, or commentary. | System prompt for the LLM that generates prompt variations. |
| variation_instructionopt | STRING | Extra instructions appended to the variation request. E.g. 'Keep it photorealistic' or 'Vary the lighting dramatically'. | |
| seedopt | INT | 00–2147483647 | Base seed. 0 = random. Each variation gets seed + variation_index. |
| max_dimensionopt | INT | 2048512–4096 | Resize source images to this max dimension before sending to API. |
| delay_secondsopt | FLOAT | 1.00–30 | Delay between API calls to avoid rate limiting. |
| harassment_thresholdopt | COMBO | BLOCK_NONE | 4 options: BLOCK_NONE, BLOCK_ONLY_HIGH, BLOCK_MEDIUM_AND_ABOVE, BLOCK_LOW_AND_ABOVE |
| hate_speech_thresholdopt | COMBO | BLOCK_NONE | 4 options: BLOCK_NONE, BLOCK_ONLY_HIGH, BLOCK_MEDIUM_AND_ABOVE, BLOCK_LOW_AND_ABOVE |
| sexually_explicit_thresholdopt | COMBO | BLOCK_NONE | 4 options: BLOCK_NONE, BLOCK_ONLY_HIGH, BLOCK_MEDIUM_AND_ABOVE, BLOCK_LOW_AND_ABOVE |
| dangerous_content_thresholdopt | COMBO | BLOCK_NONE | 4 options: BLOCK_NONE, BLOCK_ONLY_HIGH, BLOCK_MEDIUM_AND_ABOVE, BLOCK_LOW_AND_ABOVE |
Outputs (4)
| Name | Type | Description |
|---|---|---|
| images | IMAGE | — |
| log | STRING | — |
| generated_count | INT | — |
| output_folder | STRING | — |