Nodes/ComfyUI-Gemini_3x_Pro/🧠 Gemini 3.x Pro Multimodal v2
ComfyUI Node

🧠 Gemini 3.x Pro Multimodal v2

The Node That Reads Your Images, Audio, and \u201cVideo\u201d

By asirusasr-makerΒ·Created 2 months agoΒ·Updated a day agoΒ· 6
🧠 Gemini 3.x Pro Multimodal v2
  • images
  • video
  • audio
  • generated_content
  • raw_response
  • usage_info
β—„promptAnalyze this contentβ–Ί
β—„modelgemini-3.8-flashβ–Ί
β—„operation_modeanalysisβ–Ί
β—„system_instructionβ–Ί
β—„api_keyβ–Ί
β—„proxyβ–Ί
β—„temperature0.70β–Ί
β—„max_output_tokens8192β–Ί
β—„top_p0.95β–Ί
β—„top_k40β–Ί
β—„use_search_groundingfalseβ–Ί
β—„chat_modefalseβ–Ί
β—„chat_sessiondefaultβ–Ί
β—„clear_historyfalseβ–Ί
β—„fallback_enabledtrueβ–Ί
β—„retries_per_model1β–Ί
β—„cooldown_seconds30β–Ί
β—„json_schemaβ–Ί

What it actually is

This is the pack's core node, and the honest one-line version is: it puts a frontier multimodal model in the middle of your graph. Load an image, ask a question, get text back that you can wire into anything downstream. Caption a reference sheet before you feed it to a video model. Turn a sloppy sentence into a structured prompt for Z-Image or Flux. None of that is diffusion, and it doesn't touch your VRAM.

Worth being clear about what this is not: it is not the text encoder inside your checkpoint - the Qwen3 or T5 that turns your prompt into conditioning during the forward pass. That one you can't swap or skip. This is a separate model you bolt on, it runs on Google's servers, and it only runs when the node executes.

The trade is that everything you send leaves your machine and Google's filter applies. That's exactly why the local-LLM crowd runs an abliterated 8B for prompt rewriting instead. Use this node when you want frontier reading comprehension and don't mind a server seeing the image; stay local for the jobs a hosted filter would refuse.

How it works

One generate_content call per execution. The node assembles a contents list in a specific order: your images batch first (every image in the tensor), then up to 16 frames sampled from the video input, then audio re-encoded to 16 kHz mono PCM16 and sent as inline bytes, then your prompt string last. Google Search grounding, when you turn it on, is a server-side tool - nothing local to install.

The one thing to internalise: video is typed as IMAGE and is not a video upload. It samples up to 16 stills from an image tensor. The README keeps the field for workflow compatibility, and the source comment says it plainly - ComfyUI hands this input an IMAGE tensor regardless. Actual video generation is the other node.

The inputs that matter

Start with three. prompt is your instruction, model picks the Gemini 3.x tier (default gemini-3.8-flash; the seven choices run down to gemini-3.1-flash-lite, with gemini-3.1-pro-preview at the top), and operation_mode picks analysis, chat, or structured_json.

Then the plumbing: images, audio, system_instruction for the standing instruction, and the usual sampling dials (temperature, top_p, top_k, max_output_tokens). use_search_grounding lets the model search before answering. chat_mode plus chat_session gives you a conversation remembered across runs under that name, and clear_history wipes it. Be aware what that history is: an in-memory dict in the node class, keyed on session name and model, gone when ComfyUI restarts. Switch models mid-conversation and you've forked to a fresh thread.

api_key, proxy, fallback_enabled, retries_per_model and cooldown_seconds are the plumbing you'll meet on every node in this pack: the key, an endpoint override, and the retry router's three dials.

json_schema only does anything in structured_json mode, and it wants real JSON - a parse failure comes back as an error string rather than a crash.

Outputs

generated_content is the text; wire it into your prompt encoder, a Save Text node, or another generator. raw_response is the whole response object stringified, useful when you need to see what the model actually returned. usage_info is JSON, and it's the one to read when something looks wrong: requested model, actual model, whether fallback fired, session ID, token counts.

Install

ComfyUI Manager is the easy path if the registry has it - search the display name ComfyUI Gemini 3x Pro. The pack's own README documents only the manual route, which always works:

cd ComfyUI/custom_nodes
git clone https://github.com/asirusasr-maker/ComfyUI-Gemini_3x_Pro

Then install its dependencies with the Python that runs your ComfyUI - on portable Windows that's the embedded one:

python_embeded\python.exe -m pip install -r ComfyUI\custom_nodes\ComfyUI-Gemini_3x_Pro\requirements.txt

requirements.txt is short: google-genai>=2.27.0,<3.0, pillow, numpy, sounddevice. Torch is deliberately not in there, so nothing stomps your CUDA build.

Finally, the API key (Google AI Studio). Resolution order: the api_key field, then GEMINI_API_KEY in config.json, then the GEMINI_API_KEY environment variable. Restart ComfyUI.

Where people get burned

Failures come back as text, not exceptions. The node returns ("Error: ...", "", "") when something goes wrong, so the workflow completes "successfully" and hands your prompt encoder the words Error: No Gemini API key. If the output reads like an error message, it is one.

401 is yours to fix. Auth and bad-request errors aren't retried - v2 surfaces 400, 401 and 403 immediately, which is correct behaviour and also means no amount of re-queueing will help. A key left as your_api_key_here, or a wrong proxy, lands here.

429 and 503 are handled for you. Those get retried with exponential backoff plus jitter, then the node walks down the configured model list - Flash 3.8 to 3.7 to 3.6 and so on - with a 30-second cooldown on a model that keeps failing. That cooldown is in-process and resets with ComfyUI; if every candidate is cooling down, the run raises "All Gemini model candidates are temporarily unavailable."

Model IDs rot fast. Version 2 of this pack rewrote the entire catalog to match Google's October 2026 model page, and it will happen again. A workflow saved against a retired ID won't necessarily error out - a model-not-found is treated as a fallback signal, so you get quietly moved onto a different model. Read actual_model in usage_info.

CategoryGemini 3.x

Inputs (21)

NameTypeDefaultDescription
promptSTRINGAnalyze this contentβ€”
modelCOMBOgemini-3.8-flash7 options: gemini-3.8-flash, gemini-3.7-flash, gemini-3.6-flash, gemini-3.5-flash, gemini-3.5-flash-lite, gemini-3.1-flash-lite, +1
operation_modeCOMBOanalysis3 options: analysis, chat, structured_json
imagesoptIMAGEβ€”
videooptIMAGEβ€”
audiooptAUDIOβ€”
system_instructionoptSTRINGβ€”
api_keyoptSTRINGβ€”
proxyoptSTRINGβ€”
temperatureoptFLOAT0.700–1β€”
max_output_tokensoptINT81921–65536β€”
top_poptFLOAT0.950–1β€”
top_koptINT401–100β€”
use_search_groundingoptBOOLEANfalseβ€”
chat_modeoptBOOLEANfalseβ€”
chat_sessionoptSTRINGdefaultβ€”
clear_historyoptBOOLEANfalseβ€”
fallback_enabledoptBOOLEANtrueβ€”
retries_per_modeloptINT10–4β€”
cooldown_secondsoptFLOAT300–300β€”
json_schemaoptSTRINGβ€”

Outputs (3)

NameTypeDescription
generated_contentSTRINGβ€”
raw_responseSTRINGβ€”
usage_infoSTRINGβ€”