Nodes/DIGIT Nodes/DIGIT Gemini Image
ComfyUI Node

DIGIT Gemini Image

Nano Banana, in your graph, billed straight to your GCP account

By thedepartmentofexternalservices·Created 7 months ago·Updated 2 months ago· 0
DIGIT Gemini Image
  • image1
  • image2
  • image3
  • image4
  • image5
  • image6
  • image7
  • image8
  • image9
  • image
  • text
◄prompt►
◄modelgemini-3.1-flash-image►
◄aspect_ratio16:9►
◄resolution1K►
◄thinking_levelMINIMAL►
◄seed0►
◄temperature1.00►
◄gcp_project_id►
◄gcp_region►
◄system_instructionYou are an expert image-generation engine. You must ALWAYS produce an image. Interpret all user input—regardless of format, intent, or abstraction—as literal visual directives for image composition. If a prompt is conversational or lacks specific visual details, you must creatively invent a concrete visual scenario that depicts the concept. Prioritize generating the visual representation above any text, formatting, or conversational requests.►
◄top_p1.00►
◄top_k32►
◄harassment_thresholdBLOCK_NONE►
◄hate_speech_thresholdBLOCK_NONE►
◄sexually_explicit_thresholdBLOCK_NONE►
◄dangerous_content_thresholdBLOCK_NONE►
◄batch_count1►

This is the node that makes the closed frontier feel local. Gemini's image models - the family the community knows as Nano Banana - have no open weights, so if you want them, you call them. This node calls them through Google's Vertex AI, and here's the twist that sets the DIGIT pack apart: there's no API key. You authenticate once with gcloud, and every image you generate bills directly to your own GCP account at Google's list price. No proxy, no wrapper service, no markup, no third party reading your prompts.

It's a unified node, too. Prompt only → text-to-image. Prompt plus up to nine input images → edit, style transfer, or multi-image composition. Same node, auto-detected from what you plug in.

How it works

Under the hood it uses the official google-genai SDK against Vertex AI. The node figures out your project and region automatically - from the DIGIT_GCP_PROJECT / DIGIT_GCP_REGION env vars, your gcloud config, or the GCP metadata service if you're running on a Compute Engine VM. On a GCP instance you don't even set the project; it just works. On your laptop you run the one-time auth in the GCP section below.

The models available:

  • gemini-3.1-flash-image - the default. Nano Banana 2: the balanced pick.
  • gemini-3.1-flash-lite-image - Nano Banana 2 Lite. Fastest, cheapest, 1K resolution only.
  • gemini-3-pro-image - Nano Banana Pro. Higher quality, slower, pricier.
  • gemini-2.5-flash-image - the previous generation, still solid.

The inputs a beginner actually sets:

  • prompt - required, and it's a real language-model prompt. Conversational works; the node ships a system instruction that forces an image out even from abstract asks.
  • aspect_ratio and resolution - 1K/2K/4K, plus 13-ish aspect ratios. Pick the frame, not the raw pixels.
  • model - start at the default, move up when you want the Pro.
  • seed - 0 is random; set a number to reproduce a run.
  • image1…image9 - optional. Connect one and it becomes an edit; connect several and it composes. Batched images get iterated automatically.
  • batch_count - fire 1–128 generations in parallel as one IMAGE batch, each with its own seed.
  • temperature, top_p, top_k, and thinking_level (MINIMAL vs HIGH) - the sampling knobs if you want to chase consistency.

The safety thresholds (harassment, hate speech, sexual, dangerous) default to BLOCK_NONE - this is your GCP project, so Google's default filters are left alone. Leave them there unless you have a reason.

Outputs are image (an IMAGE tensor, RGBA) and text (whatever the model said alongside - usually empty).

Install and auth

cd ComfyUI/custom_nodes
git clone https://github.com/thedepartmentofexternalservices/comfyui-digit.git
cd comfyui-digit
pip install -r requirements.txt

(Or ComfyUI Manager → search comfyui-digit → install.) Restart, then the one-time GCP setup:

gcloud auth application-default login
gcloud config set project YOUR_PROJECT_ID
gcloud auth application-default set-quota-project YOUR_PROJECT_ID
gcloud services enable aiplatform.googleapis.com

That's it. No key field to paste.

Common issues

Most "why is this failing" posts in this category trace back to auth or billing. If the node throws a permissions error, you're logged in but your account can't bill that project - check that Vertex AI is enabled and your billing account is attached. If it says it can't find the project, gcloud config set project before you start ComfyUI, or set gcloud_project_id on the node.

The retry behavior is built in - automatic exponential backoff on 429 (rate limit) and 503 (service unavailable), up to three tries - so transient hiccups mostly sort themselves. What won't sort itself is quota: your GCP project has its own Vertex AI limits, and 128 parallel batch calls can hit them fast. If you're batch-generating at scale, expect to raise quota or throttle.

Two honest caveats. First, your prompts and images go to Google - that's the deal with any closed-model API, and the reason the local-only crowd skips it. Second, remember the safety thresholds default to BLOCK_NONE on your project; if you're sharing a GCP account, that's a setting someone should look at.

CategoryDIGIT

Inputs (26)

NameTypeDefaultDescription
promptSTRING—
modelCOMBOgemini-3.1-flash-image4 options: gemini-3.1-flash-image, gemini-3.1-flash-lite-image, gemini-3-pro-image, gemini-2.5-flash-image
aspect_ratioCOMBO16:913 options: auto, 1:1, 2:3, 3:2, 3:4, 4:1, +7
resolutionCOMBO1K3 options: 1K, 2K, 4K
thinking_levelCOMBOMINIMALThinking level for image generation. HIGH may improve quality.
seedINT00–2147483647—
temperatureFLOAT1.000–2—
gcp_project_idSTRINGGCP project ID. Auto-detected from DIGIT_GCP_PROJECT env var or GCP metadata.
gcp_regionSTRINGGCP region. Auto-detected from DIGIT_GCP_REGION env var or GCP metadata. Defaults to 'global'.
image1optIMAGE—
image2optIMAGE—
image3optIMAGE—
image4optIMAGE—
image5optIMAGE—
image6optIMAGE—
image7optIMAGE—
image8optIMAGE—
image9optIMAGE—
system_instructionoptSTRINGYou are an expert image-generation engine. You must ALWAYS produce an image. Interpret all user input—regardless of format, intent, or abstraction—as literal visual directives for image composition. If a prompt is conversational or lacks specific visual details, you must creatively invent a concrete visual scenario that depicts the concept. Prioritize generating the visual representation above any text, formatting, or conversational requests.—
top_poptFLOAT1.000–1—
top_koptINT321–64—
harassment_thresholdoptCOMBOBLOCK_NONE4 options: BLOCK_NONE, BLOCK_ONLY_HIGH, BLOCK_MEDIUM_AND_ABOVE, BLOCK_LOW_AND_ABOVE
hate_speech_thresholdoptCOMBOBLOCK_NONE4 options: BLOCK_NONE, BLOCK_ONLY_HIGH, BLOCK_MEDIUM_AND_ABOVE, BLOCK_LOW_AND_ABOVE
sexually_explicit_thresholdoptCOMBOBLOCK_NONE4 options: BLOCK_NONE, BLOCK_ONLY_HIGH, BLOCK_MEDIUM_AND_ABOVE, BLOCK_LOW_AND_ABOVE
dangerous_content_thresholdoptCOMBOBLOCK_NONE4 options: BLOCK_NONE, BLOCK_ONLY_HIGH, BLOCK_MEDIUM_AND_ABOVE, BLOCK_LOW_AND_ABOVE
batch_countoptINT11–128Number of images to generate. Each is a separate API call fired in parallel; results return as one IMAGE batch.

Outputs (2)

NameTypeDescription
imageIMAGE—
textSTRING—