Nodes/ComfyUI-Gemini_3x_Pro/🎬 Gemini Video Generation v2
ComfyUI Node

🎬 Gemini Video Generation v2

Eight Seconds, a Polling Loop, and the ffmpeg You Already Have

By asirusasr-makerΒ·Created 2 months agoΒ·Updated a day agoΒ· 6
🎬 Gemini Video Generation v2
  • reference_image
  • video_frames
  • video_info
  • raw_response
β—„promptA cinematic drone shot over a futuristic cityβ–Ί
β—„modelveo-3.1-generate-previewβ–Ί
β—„duration8sβ–Ί
β—„aspect_ratio16:9β–Ί
β—„resolution720pβ–Ί
β—„api_keyβ–Ί
β—„proxyβ–Ί
β—„seed-1β–Ί
β—„omni_taskautoβ–Ί
β—„fallback_enabledtrueβ–Ί
β—„retries_per_model1β–Ί
β—„cooldown_seconds45β–Ί
β—„decode_fps12β–Ί
β—„max_frames96β–Ί

The one thing open weights still can't do

Veo generates audio natively with the picture - dialogue, sound effects and ambience in the same pass as the video. That's the genuine, unclosed gap in this ecosystem: LTX-2.3 is the only open attempt at native audio-plus-video, and it's still behind. Everything else people bolt onto a silent clip in post. So if you want a talking shot that sounds like a talking shot, this node is one of the few doors, and it's a closed one.

What you're buying, though, is eight seconds. Veo isn't a scene generator; it's a shot generator. The workflows that work are the ones that treat it that way: generate establishing shots and hero beats on the API, then stitch, extend, or animate around them locally. You also can't run it without a key, and you can't run it cheaply - per-second pricing on video is where a session gets expensive fast. Going in with a shot list beats going in with a prompt.

How it works

Two different backends behind one set of inputs. For the Veo models, the node starts an async video operation and then polls it every eight seconds until the job reports done - so a single execution occupies a slot in your queue for minutes, not seconds. When the job finishes, it downloads the MP4 through the SDK's file API, writes it into ComfyUI's temp directory, and then decodes it into an IMAGE batch. Gemini Omni 1.1 Flash goes the other way, through the Interactions API, taking your prompt plus an optional image and a task hint.

That decode step is a system call, not a Python library. The node shells out to ffmpeg and ffprobe - the pack does not install either, and its requirements.txt contains no video dependency at all. Frames come out at decode_fps (default 12), capped at max_frames (default 96).

The node also protects you from your own settings when a fallback happens, and it'll do it silently. Duration and resolution get clamped to whatever the model that actually ran supports: ask for 10s and fall back to Veo and you get 8, ask for 4k on Fast and you get 1080p, ask for 1080p and fall back to Veo Lite and you get 720p. video_info records the effective values next to the requested ones.

Inputs and outputs

Required: prompt, model (four choices - the three Veo 3.1 tiers plus gemini-omni-1.1-flash), duration (4s, 6s, 8s, 10s), aspect_ratio (16:9 or 9:16), resolution (720p, 1080p, 4k).

Optional: reference_image for image-to-video - note it uses the first frame only, so a batch gets trimmed to one - seed (only sent when you set it to something other than -1, and Omni ignores it entirely), omni_task for Omni's text_to_video, image_to_video or reference_to_video modes, decode_fps and max_frames for the decode, plus api_key, proxy, fallback_enabled, retries_per_model and cooldown_seconds.

Outputs: video_frames is the decoded IMAGE batch - wire it into Save Image, a video-combine node, or an interpolator. video_info is JSON and it's the important one, because it carries video_path, the local MP4 in ComfyUI's temp folder, along with the requested-versus-effective model, duration and resolution. raw_response is the operation payload.

One thing to plan for: the soundtrack lives in that MP4, not in the frames. Save the frames and you get a silent clip. If you need the audio - and with Veo you probably do - point a video-loading node at the path from video_info rather than rebuilding the video from video_frames.

Install

Manager, searching ComfyUI Gemini 3x Pro, or the documented clone:

cd ComfyUI/custom_nodes
git clone https://github.com/asirusasr-maker/ComfyUI-Gemini_3x_Pro

Dependencies, with the interpreter that runs ComfyUI:

python_embeded\python.exe -m pip install -r ComfyUI\custom_nodes\ComfyUI-Gemini_3x_Pro\requirements.txt

Then confirm ffmpeg is visible to the ComfyUI process - being on your PATH in a terminal you opened separately doesn't count:

ffmpeg -version
ffprobe -version

Put GEMINI_API_KEY in the pack's config.json or the environment, and restart. The api_key field on the node overrides both, but it's a plain string widget, so it gets serialised into any workflow you share.

Where people get burned

Placeholder frames instead of a video. If ffmpeg or ffprobe isn't on the process's PATH, the generation still succeeds and still saves the MP4 - you just get a single 512Γ—512 black frame back instead of your clip. That looks exactly like a failed generation. Check video_info: if it has a real video_path, the video exists and only the decode failed.

It ties up the queue. The polling loop sleeps in eight-second steps while the job runs, and the generation itself takes a while. Don't queue a Veo shot last in a long batch; queue it first and go make coffee.

429 on video is common and expensive-looking. A rate limit here is the same transient class as everywhere else in the pack - retried, then fallen back down the model chain. Because a fallback can move you from Veo 3.1 to Omni and back, always check actual_model in video_info rather than assuming you got the model you picked.

Requested isn't what you got. The duration and resolution clamping above is the single most confusing thing about this node. When a clip comes back shorter or softer than asked, video_info already has the explanation in it.

CategoryGemini 3.x

Inputs (15)

NameTypeDefaultDescription
promptSTRINGA cinematic drone shot over a futuristic cityβ€”
modelCOMBOveo-3.1-generate-preview4 options: veo-3.1-generate-preview, veo-3.1-fast-generate-preview, veo-3.1-lite-generate-preview, gemini-omni-1.1-flash
durationCOMBO8s4 options: 4s, 6s, 8s, 10s
aspect_ratioCOMBO16:92 options: 16:9, 9:16
resolutionCOMBO720p3 options: 720p, 1080p, 4k
reference_imageoptIMAGEβ€”
api_keyoptSTRINGβ€”
proxyoptSTRINGβ€”
seedoptINT-1-1–2147483647β€”
omni_taskoptCOMBOauto4 options: auto, text_to_video, image_to_video, reference_to_video
fallback_enabledoptBOOLEANtrueβ€”
retries_per_modeloptINT10–3β€”
cooldown_secondsoptFLOAT450–300β€”
decode_fpsoptFLOAT121–24β€”
max_framesoptINT961–240β€”

Outputs (3)

NameTypeDescription
video_framesIMAGEβ€”
video_infoSTRINGβ€”
raw_responseSTRINGβ€”