π¨ Gemini Image Generation v2
What the Image Node Is Actually Sending
- reference_image
- generated_images
- generation_info
- raw_response
Why you'd put a cloud model in a local graph
Nano Banana - Google's Gemini-native image generation - has no open weights. You cannot download it, so calling it is the only door. What makes it worth wiring into ComfyUI anyway is everything around it: your masking, your upscaler, your LoRA, your video model, your batch logic. People generate a source frame with a closed model and then do five local operations on it, and that's a workflow ComfyUI is genuinely better at than a chat window.
The reference-image use case is where this node earns its place. One good product shot in, a batch of consistent variations out, in the same graph that already knows how to crop, resize and save them. The pattern working users describe is one source image turning into nine derivatives per iteration, without leaving ComfyUI.
Cost is per call, and it's the real trade. A local generation is free after the electricity. Accept that going in - and remember the inputs leave your machine and hit Google's moderation, however clever your graph is.
How it works
The node calls the Interactions API once per requested image, in a loop - so num_images=4 is four separate billed requests, not one request for four pictures. Each request carries input blocks: your reference image (as a base64 JPEG) if you supplied one, then a text block.
Two things about that text block are worth knowing before you spend time fighting them. negative_prompt is not a negative prompt. There's no negative conditioning on the API; the node appends your text as a line reading Avoid: <your text> onto the prompt. It's a hint, not a control, and it will sometimes get ignored. And reference_image sends exactly one image - the first frame of whatever batch you feed it. Wire a 6-image batch in and five of them are dropped on the floor.
Model names map to official IDs: Nano Banana Pro is gemini-3-pro-image, Nano Banana 2 is gemini-3.1-flash-image, Nano Banana 2 Lite is gemini-3.1-flash-lite-image. Fallback walks Pro β 2 β 2 Lite when a model is unavailable or the server is having a moment.
Inputs and outputs
Required: prompt, model (default Nano Banana 2), aspect_ratio - ten choices from 1:1 through 21:9, including the portrait 4:5 most social crops want.
Optional and actually used: reference_image, negative_prompt, num_images (1β4), image_size (1K, 2K, 4K), api_key, proxy, plus fallback_enabled, retries_per_model and cooldown_seconds for the retry router.
Outputs: generated_images is a stacked IMAGE batch of everything that came back, which is what Save Image or an upscale branch wants. generation_info is JSON - the requested and actual models, the effective image size per image, how many came back versus how many you asked for, and the prompt. raw_response is the raw API payload, handy when you're trying to work out why a prompt got a refusal.
Install
Manager, if the registry has it: search ComfyUI Gemini 3x Pro. Otherwise the pack's documented route:
cd ComfyUI/custom_nodes
git clone https://github.com/asirusasr-maker/ComfyUI-Gemini_3x_Pro
cd /path/to/your/ComfyUI
python -m pip install -r custom_nodes/ComfyUI-Gemini_3x_Pro/requirements.txt
Portable Windows users use the embedded interpreter instead:
python_embeded\python.exe -m pip install -r ComfyUI\custom_nodes\ComfyUI-Gemini_3x_Pro\requirements.txt
Grab a key from Google AI Studio and put it in the pack's config.json as GEMINI_API_KEY, or set the environment variable of the same name. The node's own api_key field overrides both, which is convenient and also a way to leak your key into a shared workflow file - string widgets are serialised into the JSON. Restart ComfyUI.
Where people get burned
A failed call returns a black image, not an error. On failure the node hands back a 512Γ512 tensor of zeros and puts the message in generation_info. Downstream that's a black frame that saves fine and upscales fine, and people chase it for an hour. Read generation_info first.
Silent fallback changes your image size. If you asked for 4K and the run quietly fell back to Lite, the node forces 1K - Lite's documented envelope tops out lower. It does the right thing technically, but you asked for 4K and got 1K with no complaint. generation_info records effective_image_sizes; check it when a render looks soft.
429 and 503 are normal at peak. Retries and model fallback handle them, with a short cooldown on a model that keeps failing. If every candidate ends up cooling down you get an error instead of an image, and the fix is patience, not a config change. 401 and 403 are the opposite - those surface immediately and will never fix themselves with a re-queue.
Model IDs move. This pack rewrote its entire model catalog in v2 to match Google's current API listing, and it'll need it again. If a workflow saved months ago starts producing from a model you didn't pick, look at actual_model in generation_info.
The filter is not your node's fault, and you can't patch it. Google's moderation applies to prompts and reference images, including the January 2026 tightening around celebrity and famous-IP generation. There are no weights to edit, so no abliterated workaround exists. If your job is the kind a hosted filter refuses, this is the wrong tool and a local checkpoint is the right one.
Inputs (12)
| Name | Type | Default | Description |
|---|---|---|---|
| prompt | STRING | A cinematic portrait | β |
| model | COMBO | Nano Banana 2 | 3 options: Nano Banana Pro, Nano Banana 2, Nano Banana 2 Lite |
| aspect_ratio | COMBO | 1:1 | 10 options: 1:1, 3:2, 2:3, 3:4, 4:3, 4:5, +4 |
| reference_imageopt | IMAGE | β | |
| negative_promptopt | STRING | β | |
| api_keyopt | STRING | β | |
| proxyopt | STRING | β | |
| num_imagesopt | INT | 11β4 | β |
| fallback_enabledopt | BOOLEAN | true | β |
| retries_per_modelopt | INT | 10β4 | β |
| cooldown_secondsopt | FLOAT | 300β300 | β |
| image_sizeopt | COMBO | 1K | 3 options: 1K, 2K, 4K |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| generated_images | IMAGE | β |
| generation_info | STRING | β |
| raw_response | STRING | β |