NanoBanana - Video Generation (Gemini Omni)
Video out of a node, and editing the clip you already have
- image
- network
- video_file_path
- interaction_id
What you're actually wiring up
Quick orientation: "NanoBanana" is a community nickname for Google's Gemini image models, and this pack blew past it - 32 nodes now. NanoBanana_OmniVideoGen calls gemini-omni-1.1-flash over Google's Interactions API and hands you back an MP4 path.
Why do that in a graph at all? Closed video models have no open weights - you can't download Omni or Veo - so an API node is the only door if you want that motion quality sitting next to your local upscaler and your masking.
The more interesting reason: Omni is unusually good at modifying footage that already exists. Community A/B tests put the split plainly - it's "excellent at modifying and editing elements that are already in the shot," and less predictable when you ask it to insert a big new object from nothing. Treat it as an editor you can wire downstream of your own rendered frames, not just a text-to-video vending machine. It also does plain text-to-video and image-to-video, with native audio in the same pass - the one capability open video still hasn't matched.
And timely: the pack's Veo node is slated for shutdown on 2026-10-22 per the README. Omni is where that traffic lands.
How it works
The implementation is short. The node builds one request and calls client.interactions.create(). Images you wire in get encoded to PNG base64 and prepended to your prompt; a Files API video URI is attached as a video part. The response format asks for video with your aspect ratio, resolution and delivery mode. Then it either base64-decodes the bytes straight into <output>/gemini_omni_<8hex>.mp4, or - with delivery set to uri - polls files.get() until the upload is ACTIVE and downloads it.
Two pack behaviors matter: calls run through retry-with-jittered-backoff, so a transient 429 doesn't kill the queue, and every node re-executes on every run by design - which means every queue is a fresh billed call.
The inputs that matter
Required are just api_key, model, and prompt. The model dropdown has exactly one entry (gemini-omni-1.1-flash) - that's the current Omni line, not an oversight. custom_model is the escape hatch for whatever Google ships next, and it's validated against a safe-ID pattern so a typo with ../ in it can't repoint the endpoint. The optional inputs are where the node earns its place:
image- the tooltip is the spec: one image is a start frame, two are first and last frame, more become subject references. The prompt is what tells the model how to use them.video_uri- a Files API URI from the pack's Files Upload node, to edit or extend footage you already have. Up to 10 seconds in.previous_interaction_id- wire theinteraction_idoutput back into this, and your prompt becomes an edit instruction on the video the earlier call produced. That loop is the point of this node.resolution-360p,720p(default),1080por4k;aspect_ratiois16:9or9:16.delivery-base64returns the video inline with roughly a 4 MB ceiling;uridownloads it through the Files API and handles bigger files.timeout_seconds- defaults to 600, tops out at 1800. A 4k pass is not fast.network- optional, for the pack's Network Route node.
Outputs are video_file_path (a STRING pointing at the MP4 in your output folder) and interaction_id. It's a string, not a video handle, so to use the clip you feed it to something that loads from a path - VideoHelperSuite's Load Video Path is the usual reach - or just open the file.
Installing it
ComfyUI Manager: search NanoBanana2 and install. Registry: comfy node registry-install nanobanana2. Manual:
cd ComfyUI/custom_nodes
git clone https://github.com/IxMxAMAR/ComfyUI-NanoBanana2
cd ComfyUI-NanoBanana2
pip install "google-genai>=2.3.0"
Restart after. You need a Google AI Studio key (aistudio.google.com → Get API Key); paste it in the password-masked API Key field, or set GEMINI_API_KEY and leave the field blank.
The google-genai >= 2.3.0 floor is real, not decorative - it's the first release with the output_video helper this node reads. One mismatch worth knowing: requirements.txt lists only google-genai, while pyproject.toml also asks for httpx[socks]. If you installed via Manager and later try the Network Route node against a SOCKS proxy, that's the extra you're missing.
Where people get burned
- Cost. Video is where a session gets expensive fast, and this is metered per call. Price a batch before you run one.
- Filtering follows the model, not the node. It's a closed Google model; whatever it refuses, this refuses, and there are no weights to abliterate. That's the exact axis where a local model wins.
- Your inputs leave the machine. Prompt, reference images, uploaded video - all of it goes to Google. That's the mechanism, not a bug. The base64 ~4 MB ceiling is the classic first-4k surprise too; if a big render comes back truncated, you wanted
uri. - "Gemini Omni returned no video." That means the call completed but the response carried no video part. The console prints which model was used right before it; re-read your prompt and try again.
If your shot needs a brand-new element locked to an exact camera move, this isn't the pick. For editing what's already in frame, feeding it your own frames, and iterating quickly, it's a solid fit inside the graph.
Inputs (12)
| Name | Type | Default | Description |
|---|---|---|---|
| api_key | STRING | — | |
| model | COMBO | gemini-omni-1.1-flash | 1 options: gemini-omni-1.1-flash |
| prompt | STRING | — | |
| custom_modelopt | STRING | — | |
| imageopt | IMAGE | Optional image(s): one = start frame, two = first and last frame, more = subject references. The prompt says how to use them. | |
| video_uriopt | STRING | Files API URI of an uploaded video (Files Upload node) to edit or extend. Inputs up to 10 seconds. | |
| previous_interaction_idopt | STRING | interaction_id of an earlier Omni call; the prompt then edits that video. | |
| aspect_ratioopt | COMBO | 16:9 | 2 options: 16:9, 9:16 |
| resolutionopt | COMBO | 720p | 4 options: 720p, 360p, 1080p, 4k |
| deliveryopt | COMBO | base64 | base64 returns the video inline (about 4 MB limit). uri downloads it through the Files API and supports larger videos. |
| timeout_secondsopt | INT | 60060–1800 | Max time to wait for the video. |
| networkopt | NB_NETWORK | Optional. Wire a NanoBanana - Network Route node here to route this request through that proxy (e.g. US egress). |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| video_file_path | STRING | — |
| interaction_id | STRING | — |