DIGIT Gemini Image
Nano Banana, in your graph, billed straight to your GCP account
- image1
- image2
- image3
- image4
- image5
- image6
- image7
- image8
- image9
- image
- text
This is the node that makes the closed frontier feel local. Gemini's image models - the family the community knows as Nano Banana - have no open weights, so if you want them, you call them. This node calls them through Google's Vertex AI, and here's the twist that sets the DIGIT pack apart: there's no API key. You authenticate once with gcloud, and every image you generate bills directly to your own GCP account at Google's list price. No proxy, no wrapper service, no markup, no third party reading your prompts.
It's a unified node, too. Prompt only → text-to-image. Prompt plus up to nine input images → edit, style transfer, or multi-image composition. Same node, auto-detected from what you plug in.
How it works
Under the hood it uses the official google-genai SDK against Vertex AI. The node figures out your project and region automatically - from the DIGIT_GCP_PROJECT / DIGIT_GCP_REGION env vars, your gcloud config, or the GCP metadata service if you're running on a Compute Engine VM. On a GCP instance you don't even set the project; it just works. On your laptop you run the one-time auth in the GCP section below.
The models available:
gemini-3.1-flash-image- the default. Nano Banana 2: the balanced pick.gemini-3.1-flash-lite-image- Nano Banana 2 Lite. Fastest, cheapest, 1K resolution only.gemini-3-pro-image- Nano Banana Pro. Higher quality, slower, pricier.gemini-2.5-flash-image- the previous generation, still solid.
The inputs a beginner actually sets:
- prompt - required, and it's a real language-model prompt. Conversational works; the node ships a system instruction that forces an image out even from abstract asks.
- aspect_ratio and resolution - 1K/2K/4K, plus 13-ish aspect ratios. Pick the frame, not the raw pixels.
- model - start at the default, move up when you want the Pro.
- seed - 0 is random; set a number to reproduce a run.
- image1…image9 - optional. Connect one and it becomes an edit; connect several and it composes. Batched images get iterated automatically.
- batch_count - fire 1–128 generations in parallel as one IMAGE batch, each with its own seed.
- temperature, top_p, top_k, and thinking_level (MINIMAL vs HIGH) - the sampling knobs if you want to chase consistency.
The safety thresholds (harassment, hate speech, sexual, dangerous) default to BLOCK_NONE - this is your GCP project, so Google's default filters are left alone. Leave them there unless you have a reason.
Outputs are image (an IMAGE tensor, RGBA) and text (whatever the model said alongside - usually empty).
Install and auth
cd ComfyUI/custom_nodes
git clone https://github.com/thedepartmentofexternalservices/comfyui-digit.git
cd comfyui-digit
pip install -r requirements.txt
(Or ComfyUI Manager → search comfyui-digit → install.) Restart, then the one-time GCP setup:
gcloud auth application-default login
gcloud config set project YOUR_PROJECT_ID
gcloud auth application-default set-quota-project YOUR_PROJECT_ID
gcloud services enable aiplatform.googleapis.com
That's it. No key field to paste.
Common issues
Most "why is this failing" posts in this category trace back to auth or billing. If the node throws a permissions error, you're logged in but your account can't bill that project - check that Vertex AI is enabled and your billing account is attached. If it says it can't find the project, gcloud config set project before you start ComfyUI, or set gcloud_project_id on the node.
The retry behavior is built in - automatic exponential backoff on 429 (rate limit) and 503 (service unavailable), up to three tries - so transient hiccups mostly sort themselves. What won't sort itself is quota: your GCP project has its own Vertex AI limits, and 128 parallel batch calls can hit them fast. If you're batch-generating at scale, expect to raise quota or throttle.
Two honest caveats. First, your prompts and images go to Google - that's the deal with any closed-model API, and the reason the local-only crowd skips it. Second, remember the safety thresholds default to BLOCK_NONE on your project; if you're sharing a GCP account, that's a setting someone should look at.
Inputs (26)
| Name | Type | Default | Description |
|---|---|---|---|
| prompt | STRING | — | |
| model | COMBO | gemini-3.1-flash-image | 4 options: gemini-3.1-flash-image, gemini-3.1-flash-lite-image, gemini-3-pro-image, gemini-2.5-flash-image |
| aspect_ratio | COMBO | 16:9 | 13 options: auto, 1:1, 2:3, 3:2, 3:4, 4:1, +7 |
| resolution | COMBO | 1K | 3 options: 1K, 2K, 4K |
| thinking_level | COMBO | MINIMAL | Thinking level for image generation. HIGH may improve quality. |
| seed | INT | 00–2147483647 | — |
| temperature | FLOAT | 1.000–2 | — |
| gcp_project_id | STRING | GCP project ID. Auto-detected from DIGIT_GCP_PROJECT env var or GCP metadata. | |
| gcp_region | STRING | GCP region. Auto-detected from DIGIT_GCP_REGION env var or GCP metadata. Defaults to 'global'. | |
| image1opt | IMAGE | — | |
| image2opt | IMAGE | — | |
| image3opt | IMAGE | — | |
| image4opt | IMAGE | — | |
| image5opt | IMAGE | — | |
| image6opt | IMAGE | — | |
| image7opt | IMAGE | — | |
| image8opt | IMAGE | — | |
| image9opt | IMAGE | — | |
| system_instructionopt | STRING | You are an expert image-generation engine. You must ALWAYS produce an image. Interpret all user input—regardless of format, intent, or abstraction—as literal visual directives for image composition. If a prompt is conversational or lacks specific visual details, you must creatively invent a concrete visual scenario that depicts the concept. Prioritize generating the visual representation above any text, formatting, or conversational requests. | — |
| top_popt | FLOAT | 1.000–1 | — |
| top_kopt | INT | 321–64 | — |
| harassment_thresholdopt | COMBO | BLOCK_NONE | 4 options: BLOCK_NONE, BLOCK_ONLY_HIGH, BLOCK_MEDIUM_AND_ABOVE, BLOCK_LOW_AND_ABOVE |
| hate_speech_thresholdopt | COMBO | BLOCK_NONE | 4 options: BLOCK_NONE, BLOCK_ONLY_HIGH, BLOCK_MEDIUM_AND_ABOVE, BLOCK_LOW_AND_ABOVE |
| sexually_explicit_thresholdopt | COMBO | BLOCK_NONE | 4 options: BLOCK_NONE, BLOCK_ONLY_HIGH, BLOCK_MEDIUM_AND_ABOVE, BLOCK_LOW_AND_ABOVE |
| dangerous_content_thresholdopt | COMBO | BLOCK_NONE | 4 options: BLOCK_NONE, BLOCK_ONLY_HIGH, BLOCK_MEDIUM_AND_ABOVE, BLOCK_LOW_AND_ABOVE |
| batch_countopt | INT | 11–128 | Number of images to generate. Each is a separate API call fired in parallel; results return as one IMAGE batch. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| image | IMAGE | — |
| text | STRING | — |