Nodes/Banana Studio/Banana Studio
ComfyUI Node

Banana Studio

Banana Studio calls Gemini image models with your own key

By tjcccc·Created 10 months ago·Updated 5 months ago· 1
Banana Studio
  • image_1
  • image_2
  • image_3
  • image_4
  • image_5
  • image_6
  • images
  • logs
◄api_key►
◄modelgemini-3-pro-image-preview►
◄promptHello, Banana Studio,►
◄batch_size1►
◄aspect_ratioAuto►
◄resolution1K►
◄temperature0.50►
◄top_p0.95►
◄thinking_levelMinimal►
◄seed-1►
◄proxy►

If you've been anywhere near AI image generation this year, you've heard of Nano Banana. That's Google's branding for its Gemini image models - Nano Banana Pro (Gemini 3 Pro Image) is the 4K flagship, Nano Banana 2 (Gemini 3.1 Flash) is the speed hybrid, and the original Nano Banana (Gemini 2.5 Flash Image) is the consumer base. They're closed, API-only, and unnervingly good at instruction-following and in-image text, which is why so many people want them inside their ComfyUI workflows. Banana Studio is that door: it calls Google's Gemini API directly with your own key and drops the result back into the graph as a normal IMAGE tensor. No Comfy partner-node storefront, no credits, no Comfy account - just a Gemini API key.

The name does two jobs. "Banana" is the nod to Nano Banana; "Studio" is the umbrella for the pack's prompt utilities. The generator itself is the BananaStudio node, and it's the only node in the pack that needs anything from you beyond an API key.

What it actually does

Hit Queue and it POSTs your prompt (plus up to six images, base64-encoded as PNGs, if you've wired any) to generativelanguage.googleapis.com/v1beta/models/{model}:generateContent, with the key riding in the request header. The response carries base64 image data back, which the node decodes into a tensor for the rest of your graph. From the canvas it looks like any other generator node. The honest difference from a local model: every call is metered and your prompt leaves your machine, subject to Google's logging and its filters. That's the standing trade of this whole category, not a bug in this node.

The inputs that matter

The model dropdown holds the three Gemini image models, defaulting to gemini-3-pro-image-preview. prompt is your text instruction - Gemini reads natural language, so write like you're briefing a photographer rather than stacking tags. batch_size (1–8) generates several images per run; with a random seed each item gets its own seed (base seed plus index), and a failed item doesn't kill the rest of the batch. image_1 through image_6 are optional reference images - the API version of a character sheet, which is the pattern people lean on for consistency across generations.

aspect_ratio (Auto, 1:1, 9:16, 16:9, 3:4, 4:3, 3:2, 2:3, 5:4, 4:5, 21:9, plus the extreme 4:1/1:4/8:1/1:8 that the Flash model reserves) and resolution (512/1K/2K/4K; 512 is Flash-only) set the output format. temperature, top_p, thinking_level (Minimal/High, Flash-only), and seed are the sampling controls - and yes, a minus temperature disables it, per the node's own tooltip.

Getting the key in

You need a Google Gemini API key. You can paste one into the api_key widget for a quick test, but the proper setup is a config.ini in the pack's folder:

[auth]
GEMINI_API_KEY = your_key_here

Resolution order is config.ini first, then the node widget, then an error if neither exists. The repo git-ignores config.ini and resolves the key at runtime, so a saved workflow carries the key's name but never the secret. One honest warning straight from the README: set proxy only when you truly need it - large inline image uploads are the first thing to time out through a flaky proxy.

Outputs

images is the tensor (or batched tensor) you feed into Save Image, Preview, or any post-processing you have wired. logs is a text summary with generation status and token usage - handy if you want to actually watch your spend.

Where people get burned

This is a closed Google model, so the censorship is welded into the weights. If the model's IMAGE_SAFETY decides your prompt is a refusal case, you get no image back and the node raises "No images returned … (likely blocked by safety policy)". ComfyUI's own lead dev has said it plainly about this family of models: it's not the node's fault, Google safetymaxxed the model. Every output also carries an invisible SynthID watermark. And it's metered - Nano Banana Pro runs in the ballpark of $0.04–0.24 per image depending on resolution. It's a genuinely good node to have in the graph. It is not a replacement for your local stack, and it never pretends to be.

Install

Standard pack install, no model downloads:

cd ComfyUI/custom_nodes
git clone https://github.com/tjcccc/comfyui_banana_studio.git

Then restart ComfyUI - or search "Banana Studio" in ComfyUI Manager. The pack declares zero pip dependencies (it uses requests, torch, and PIL, all already in a stock ComfyUI), requires Python 3.10+, and needs ComfyUI ≥ 0.19.3.

CategoryBanana Studio

Inputs (17)

NameTypeDefaultDescription
api_keySTRING—
modelCOMBOgemini-3-pro-image-preview3 options: gemini-3.1-flash-image-preview, gemini-3-pro-image-preview, gemini-2.5-flash-image
promptSTRINGHello, Banana Studio,—
batch_sizeINT11–8—
aspect_ratioCOMBOAutoOnly for gemini-3-pro-image-preview. 4:1, 1:4, 8:1, and 1:8 are only for gemini-3.1-flash-image-preview.
resolutionoptCOMBO1KOnly for gemini-3-pro-image-preview. 512 is only for gemini-3.1-flash-image-preview.
temperatureoptFLOAT0.50-1–1The temperature range from 0 to 1. Set minus value to disable it.
top_poptFLOAT0.950–1—
thinking_leveloptCOMBOMinimalOnly for gemini-3.1-flash-image-preview. It controls how much the model thinks before generating images. High thinking level may lead to better quality but longer generation time.
seedoptINT-1-1–102400—
image_1optIMAGE—
image_2optIMAGE—
image_3optIMAGE—
image_4optIMAGE—
image_5optIMAGE—
image_6optIMAGE—
proxyoptSTRING—

Outputs (2)

NameTypeDescription
imagesIMAGE—
logsSTRING—