Nodes/DIGIT Nodes/DIGIT Batch Gemini Image
ComfyUI Node

DIGIT Batch Gemini Image

Batch Gemini image generation with an LLM writing the prompt variations for you

By thedepartmentofexternalservices·Created 7 months ago·Updated 2 months ago· 0
DIGIT Batch Gemini Image
    • images
    • log
    • generated_count
    • output_folder
    ◄image_folder►
    ◄prompt►
    ◄variations_per_image3►
    ◄image_modelgemini-3.1-flash-image►
    ◄llm_modelgemini-2.5-flash►
    ◄aspect_ratio16:9►
    ◄resolution1K►
    ◄thinking_levelMINIMAL►
    ◄temperature1.00►
    ◄variation_temperature1.00►
    ◄gcp_project_id►
    ◄gcp_region►
    ◄output_subfoldergemini_output►
    ◄system_instructionYou are an expert image-generation engine. You must ALWAYS produce an image. Interpret all user input—regardless of format, intent, or abstraction—as literal visual directives for image composition. If a prompt is conversational or lacks specific visual details, you must creatively invent a concrete visual scenario that depicts the concept. Prioritize generating the visual representation above any text, formatting, or conversational requests.►
    ◄variation_system_promptYou are a creative prompt engineer for image generation. Given a base prompt, produce a variation that preserves the core intent and subject but introduces creative differences in style, mood, lighting, composition, color palette, or artistic approach. Output ONLY the varied prompt text — no preamble, numbering, or commentary.►
    ◄variation_instruction►
    ◄seed0►
    ◄max_dimension2048►
    ◄delay_seconds1.0►
    ◄harassment_thresholdBLOCK_NONE►
    ◄hate_speech_thresholdBLOCK_NONE►
    ◄sexually_explicit_thresholdBLOCK_NONE►
    ◄dangerous_content_thresholdBLOCK_NONE►

    Most image-generation nodes make you write the prompt. DIGIT Batch Gemini Image is the one that skips the boring part: it takes a folder of source images, has a Gemini LLM invent prompt variations for you, and generates multiple Gemini images per source image, saving everything to an output subfolder. If you've ever needed fifty takes on the same subject - a product in different scenes, a character in different lighting, a style spread for a client - this is the node that gets you there without you typing fifty prompts.

    It's a cloud node, so the tradeoff that applies to all of them applies here too: no local GPU needed, your images leave the machine, and every generation is a metered API call through Google Vertex AI. The payoff is that Nano Banana–class image generation (Gemini 3.1 Flash Image is the default) is genuinely good, and you're not spending your evening babysitting a queue.

    How it works

    The flow has two LLM stages. First, for each source image, an LLM (llm_model, default gemini-2.5-flash) takes your base prompt and writes variations_per_image creative variations - same core subject, different style, mood, lighting, composition. Then the image model (image_model, default gemini-3.1-flash-image) generates one image per variation. The source image itself is sent along as the visual reference, resized to max_dimension (2048) first to keep the calls cheap.

    Inputs that actually matter:

    • image_folder - where your source images live. Required.
    • prompt - your base prompt; the LLM varies this per image.
    • variations_per_image - default 3, up to 50. This is your cost multiplier, so watch it.
    • image_model / llm_model - the two model dropdowns. Fast image models and a cheap LLM (flash) keep this affordable.
    • aspect_ratio, resolution - 13 ratios, and 1K/2K/4K. Note the Lite image model is 1K-only.
    • variation_temperature - higher means wilder variations. Crank it when your first batch comes back too samey.
    • thinking_level - MINIMAL (default) or HIGH. HIGH can improve quality but costs more and takes longer.

    There are also per-category safety thresholds (harassment, hate speech, sexual, dangerous), all defaulting to BLOCK_NONE - a deliberate choice that surprises people, so flip them if you want Google's filters on. seed gives reproducible runs (each variation gets seed + index). output_subfolder (default gemini_output) is created inside image_folder automatically.

    Outputs: images (the generated IMAGE batch, straight into your graph), log, generated_count, and output_folder so you know where everything landed.

    Installing and wiring it

    It's part of the digit-comfyui pack - ComfyUI Manager, search comfyui-digit, install, restart. Or:

    cd ComfyUI/custom_nodes
    git clone https://github.com/thedepartmentofexternalservices/comfyui-digit.git
    cd comfyui-digit
    pip install -r requirements.txt
    

    Authentication is the shared pack setup: gcloud with gcloud auth application-default login, a project with the Vertex AI API enabled, and either a gcp_project_id on the node, the DIGIT_GCP_PROJECT env var, or running on a GCP instance where it's auto-detected.

    Gotchas

    The big one is cost control - variations_per_image times your batch is how the bill grows, and the live cost is per generation. Start with 2–3 variations and one or two source images to sanity-check quality before you launch the full folder. Also, this is Gemini's own filter running underneath, so the "BLOCK_NONE" thresholds won't make Google's model do things its base refuses - the node just isn't adding extra moderation on top. If a generation comes back with something odd, check log for the per-image status rather than guessing.

    CategoryDIGIT

    Inputs (23)

    NameTypeDefaultDescription
    image_folderSTRINGPath to folder containing source images.
    promptSTRINGBase prompt for Gemini image generation. The LLM will create variations of this.
    variations_per_imageINT31–50How many prompt variations (and image generations) per source image.
    image_modelCOMBOgemini-3.1-flash-image4 options: gemini-3.1-flash-image, gemini-3.1-flash-lite-image, gemini-3-pro-image, gemini-2.5-flash-image
    llm_modelCOMBOgemini-2.5-flashGemini model used to generate prompt variations.
    aspect_ratioCOMBO16:913 options: auto, 1:1, 2:3, 3:2, 3:4, 4:1, +7
    resolutionCOMBO1K3 options: 1K, 2K, 4K
    thinking_levelCOMBOMINIMALThinking level for image generation. HIGH may improve quality.
    temperatureFLOAT1.000–2Temperature for image generation.
    variation_temperatureFLOAT1.000–2Temperature for LLM prompt variation. Higher = more creative variations.
    gcp_project_idSTRINGGCP project ID.
    gcp_regionSTRINGGCP region.
    output_subfolderoptSTRINGgemini_outputSubfolder name inside image_folder for outputs. Created automatically.
    system_instructionoptSTRINGYou are an expert image-generation engine. You must ALWAYS produce an image. Interpret all user input—regardless of format, intent, or abstraction—as literal visual directives for image composition. If a prompt is conversational or lacks specific visual details, you must creatively invent a concrete visual scenario that depicts the concept. Prioritize generating the visual representation above any text, formatting, or conversational requests.System instruction for Gemini image generation.
    variation_system_promptoptSTRINGYou are a creative prompt engineer for image generation. Given a base prompt, produce a variation that preserves the core intent and subject but introduces creative differences in style, mood, lighting, composition, color palette, or artistic approach. Output ONLY the varied prompt text — no preamble, numbering, or commentary.System prompt for the LLM that generates prompt variations.
    variation_instructionoptSTRINGExtra instructions appended to the variation request. E.g. 'Keep it photorealistic' or 'Vary the lighting dramatically'.
    seedoptINT00–2147483647Base seed. 0 = random. Each variation gets seed + variation_index.
    max_dimensionoptINT2048512–4096Resize source images to this max dimension before sending to API.
    delay_secondsoptFLOAT1.00–30Delay between API calls to avoid rate limiting.
    harassment_thresholdoptCOMBOBLOCK_NONE4 options: BLOCK_NONE, BLOCK_ONLY_HIGH, BLOCK_MEDIUM_AND_ABOVE, BLOCK_LOW_AND_ABOVE
    hate_speech_thresholdoptCOMBOBLOCK_NONE4 options: BLOCK_NONE, BLOCK_ONLY_HIGH, BLOCK_MEDIUM_AND_ABOVE, BLOCK_LOW_AND_ABOVE
    sexually_explicit_thresholdoptCOMBOBLOCK_NONE4 options: BLOCK_NONE, BLOCK_ONLY_HIGH, BLOCK_MEDIUM_AND_ABOVE, BLOCK_LOW_AND_ABOVE
    dangerous_content_thresholdoptCOMBOBLOCK_NONE4 options: BLOCK_NONE, BLOCK_ONLY_HIGH, BLOCK_MEDIUM_AND_ABOVE, BLOCK_LOW_AND_ABOVE

    Outputs (4)

    NameTypeDescription
    imagesIMAGE—
    logSTRING—
    generated_countINT—
    output_folderSTRING—