Nodes/ComfyUI-Nanobanana/ComfyUI-Nanobanana
ComfyUI Node

ComfyUI-Nanobanana

Run Gemini image generation inside your graph, your way

By yvtouyvtou·Created 3 months ago·Updated 3 months ago· 1
ComfyUI-Nanobanana
  • image1
  • image2
  • image3
  • image4
  • image5
  • image6
  • image7
  • image8
  • image9
  • image10
  • image11
  • image12
  • image13
  • image14
  • image
  • response
  • image_url
  • raw_response
◄protocolgemini_native►
◄base_urlhttps://your-provider.example.com►
◄api_key►
◄modelgemini-3.1-flash-image►
◄prompt—►
◄modeauto►
◄aspect_ratioauto►
◄image_size2K►
◄openai_api_stylechat_completions►
◄auth_modeauto►
◄timeout_seconds600►
◄seed0►
◄endpoint_path►
◄response_formatauto►
◄extra_body_json{}►

The name does a lot of work. ComfyUI-Nanobanana (class LinkAPIGeminiImage, in the small yvtouyvtou/ComfyUI-Yvtou-Gemini pack) is not a Nano Banana node in the way you're hoping: no model ships with it, it isn't Google's official Partner Node, and it doesn't even need a Google account. It's a configurable HTTP wrapper - send a prompt and images to any Gemini image API, and the result drops back into the graph as a normal IMAGE tensor. Point it at Google, a reseller, or an OpenAI-compatible gateway; same node, same workflow.

That flexibility is the whole pitch. The official ComfyUI API Nodes are a metered storefront over closed models like Nano Banana Pro - prepaid credits, Comfy's vetting. This is the DIY alternative: your key, your base_url, whatever endpoint you give it, for cost, region, or billing.

How it works

Underneath it's a plain requests client - the pack's only dependency (requests>=2.28.0), no model files, no weights. You feed it a prompt (and optionally up to 14 input images, mirroring Nano Banana Pro's 14-reference ceiling), it base64-encodes your tensors, POSTs to the endpoint, parses the response, and decodes the images back into a ComfyUI batch.

Two protocols. gemini_native (default) hits Gemini's /v1beta/models/{model}:generateContent endpoint; openai_compatible targets OpenAI-shaped routes instead (pick chat_completions or images_api via openai_api_style). The default model is gemini-3.1-flash-image - that's Nano Banana 2.

Output normalization is the thoughtful part. If you wire in image1, every returned image is center-cropped and scaled to image1's resolution - no black bars, no stretching - so the result is always a legal batch tensor for ComfyUI. If no image1 is connected, it uses the largest returned image instead.

The inputs that matter

Nineteen inputs; you'll touch five:

  • protocol - gemini_native or openai_compatible.
  • base_url - the one you must not skip; the default is literally the placeholder https://your-provider.example.com, and nothing works until you change it. For Google use https://generativelanguage.googleapis.com.
  • api_key - leave it empty and use env vars if you can (see below).
  • model - model name as your provider lists it.
  • prompt - plus mode (auto / text2img / img2img): auto edits if any image is connected, generates otherwise. aspect_ratio, image_size, and seed cover the Gemini path; extra_body_json is the escape hatch for provider-specific fields like {"size": "1792x1024", "quality": "hd"}.

Outputs: image (the tensor - wire it into a Save Image or VAE decode), plus response (metadata, any text the model returned), image_url, and raw_response (redacted JSON) if you want to peek.

Installing it

ComfyUI Manager is the easy path once the pack lands on the registry (search "ComfyUI-Nanobanana" or "yvtou-gemini"); until then:

cd ComfyUI/custom_nodes
git clone https://github.com/yvtouyvtou/ComfyUI-Yvtou-Gemini.git
cd ../..
python -m pip install -r custom_nodes/ComfyUI-Yvtou-Gemini/requirements.txt

Restart ComfyUI and the node appears under api/Gemini. No model downloads, no VRAM cost - just a couple of Python files and a requests install.

Where people get burned

The key in your workflow file. The README says it straight: a key typed into the api_key widget gets saved into your workflow JSON - a leak waiting to happen if you share workflows. The node resolves keys widget → YVTOU_GEMINI_API_KEY → GEMINI_API_KEY → OPENAI_API_KEY, so set an env var and leave the widget blank.

The filters follow the model. Google's moderation applies no matter which gateway you point at. The Nano Banana line is among the most aggressively filtered in the business - a January 2026 update tightened famous-IP and celebrity restrictions - and neither a node setting nor a reseller's "looser filtering" claim changes what the model refuses. Trip IMAGE_SAFETY and the fix is a local model, not a config tweak.

Error codes. The usual suspects: 400 means your params are off (check model name, extra_body_json), 401/403 means auth (key, or the wrong auth_mode - Google prefers x_goog_api_key, which keeps the key out of the URL), 404 means base_url or endpoint_path is wrong. Timeouts happen on slow gateways - bump timeout_seconds (default 600, max 3600). And every call is metered: it adds up faster than you'd budget for.

One last thing: this is a fresh, small pack with no community footprint, so read the source before trusting it with a real key - the category has been weaponized once before, and "paste your key here, it just works" is exactly the shape to inspect first. This one is MIT-licensed, a few hundred lines, and the redaction of keys and base64 blobs from raw_response suggests an author who thought about the attack surface. Good sign, not a guarantee.

Categoryapi/Gemini

Inputs (29)

NameTypeDefaultDescription
protocolCOMBOgemini_native2 options: gemini_native, openai_compatible
base_urlSTRINGhttps://your-provider.example.com—
api_keySTRING—
modelSTRINGgemini-3.1-flash-image—
promptSTRING—
modeCOMBOauto3 options: auto, text2img, img2img
aspect_ratioCOMBOauto15 options: auto, 1:1, 16:9, 9:16, 4:3, 3:4, +9
image_sizeCOMBO2K5 options: auto, 512, 1K, 2K, 4K
openai_api_styleCOMBOchat_completions2 options: chat_completions, images_api
auth_modeCOMBOauto4 options: auto, query_key, bearer, x_goog_api_key
timeout_secondsINT60010–3600—
seedINT00–2147483647—
endpoint_pathSTRING—
response_formatCOMBOauto3 options: auto, b64_json, url
extra_body_jsonSTRING{}—
image1optIMAGE—
image2optIMAGE—
image3optIMAGE—
image4optIMAGE—
image5optIMAGE—
image6optIMAGE—
image7optIMAGE—
image8optIMAGE—
image9optIMAGE—
image10optIMAGE—
image11optIMAGE—
image12optIMAGE—
image13optIMAGE—
image14optIMAGE—

Outputs (4)

NameTypeDescription
imageIMAGE—
responseSTRING—
image_urlSTRING—
raw_responseSTRING—