Nodes/Zura Nodes/Zura V4 · Generative scene-aware multicam (experimental)
ComfyUI Node

Zura V4 · Generative scene-aware multicam (experimental)

Generative Multicam That Moves the Set

By ZURAVFX·Created 22 days ago·Updated about 16 hours ago· 1
Zura V4 · Generative scene-aware multicam (experimental)
  • source_frames
  • source_audio
  • stage
  • reference_video
  • moge_geometry
  • model
  • clip
  • video_vae
  • audio_vae
  • scene_video
  • pose_video
  • model_patch
  • multicam_frames
  • original_audio
◄shot_plan—►
◄seed42►
◄steps8►
◄promptcrossview. <Picture 1> defines only the selected camera angle. Keep the exact performer, face, clothes, gestures and speaking performance from <Video 1>, with lips faithful to its original speech. Do not replace the performer.►
◄look_enabledfalse►
◄render_selected_intervals_onlyfalse►
◄use_camera_warp_guidetrue►
◄background_modeKeep source background►
◄pose_strength0.45►
◄pose_end_percent0.55►

Zura V4 · Generative scene-aware multicam (experimental) is the pack's answer to a real limitation of its own V3 renderer: V3 reproduces your camera angles but keeps your room. V4 throws the room away and builds a new one, then shoots the performer in it from each planned angle.

The "(experimental)" in the display name is the author's, and it's earned. This is the branch where the scene, the pose and the identity all have to hold at the same time.

What it adds over V3

Same required inputs - source_frames, source_audio, shot_plan, stage, seed, steps, prompt - and the same outputs, multicam_frames and original_audio. It's a subclass of the V3 node, so everything you already learned applies. Three additions:

  • pose_video - a pose-control guide, length-matched to the source.
  • model_patch - the patch the pose control applies through.
  • pose_strength (default 0.45) and pose_end_percent (default 0.55). The tooltips are the author being useful: lower the strength if pose skeleton lines show up in the output, and stop pose control partway through sampling to let the scene clean up. That second one is the tell that this is a real production recipe, not a demo - you're letting structure guide the early denoising and then letting the model paint.

How scene mode engages

Two widgets arm it: look_enabled on, and background_mode set to anything other than "Keep source background". Once armed, the node changes what it asks you to connect - moge_geometry and the camera warp are no longer used, and it wants reference_video, pose_video, model_patch, model, clip, video_vae and audio_vae instead.

And it gets strict about the plan: every shot must use a generated angle. An original-camera shot reintroduces the source room by definition, so the node raises Choose a generated angle for every V4 scene-swap shot; Original would reintroduce the source room. That's a correct refusal.

The mechanism, and why it isn't just a prompt

Instead of warping the reference video to a new viewpoint, scene mode generates from a relit start guide: the saved angle still is used as both the first and last keyframe of an H3 image-to-video conditioning, so the shot is pinned to the same physical set at both ends and can't wander. The prompt gets a matching paragraph - the first and last keyframes define the same physical set and relit performer, stay in this exact environment for the whole shot, follow the supplied pose control and the original speech.

Then pose control is applied through MiniMaxH3FunControlNetApply against your pose_video, running from 0% to pose_end_percent at pose_strength. That's how the performer's actual gestures survive the scene change - the new room is generated, the body isn't.

Same as V3, the 17k+5 H3 length grid applies, and guides get their last frame repeated so the conditioning isn't silently cropped.

Why the parallax problem is the hard part

Here's the honest engineering read: the whole reason V3 used a MoGe-based warp is that single-image models can't invent what's behind an occluder. A depth map describes one viewpoint; move the camera and you need pixels that were never captured, which something has to hallucinate. V4 leans on generated keyframes and pose rather than warped geometry, which sidesteps some of that - but it means the room's consistency is now the model's job across every shot. Expect the set to drift between angles, which is exactly what the V4 orbit prompt node is trying to argue it out of.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/ZURAVFX/ComfyUI_zura_nodes
pip install -r ComfyUI_zura_nodes/requirements.txt

You need the same ecosystem as V3 - MiniMax H3 support nodes, CrossViewWarp, KJNodes, VideoHelperSuite - plus whatever produces your pose_video and your model_patch. No weights ship with the pack. And the same licence caveat as the whole H3 path: the MiniMax H3 community licence excludes the US, EU, UK and South Korea from running the local weights.

The workflow around it

Don't run this cold. The pack's own guidance is to plan a short clip and do a Plan/Render review before processing anything long, and here that's essential: identity, room consistency and lip sync are three different failure modes in one node. Also remember that the pack's companion nodes exist for a reason - Zura · Person Motion Control makes the foreground plate, Zura · Match Mouth Motion fixes the guide's mouth timing, and Zura Klein Look Presets supplies the relight wording. V4 is the last node in that chain, not the only one.

CategoryZura/video

Inputs (22)

NameTypeDefaultDescription
source_framesIMAGE—
source_audioAUDIO—
shot_planSTRING—
stageZURA_MULTICAM_STAGE—
seedINT420–18446744073709550000—
stepsINT81–60—
promptSTRINGcrossview. <Picture 1> defines only the selected camera angle. Keep the exact performer, face, clothes, gestures and speaking performance from <Video 1>, with lips faithful to its original speech. Do not replace the performer.—
reference_videooptIMAGE—
moge_geometryoptMOGE_GEOMETRY—
modeloptMODEL—
clipoptCLIP—
video_vaeoptVAE—
audio_vaeoptVAE—
scene_videooptIMAGE—
look_enabledoptBOOLEANfalse—
render_selected_intervals_onlyoptBOOLEANfalse—
use_camera_warp_guideoptBOOLEANtrue—
background_modeoptSTRINGKeep source background—
pose_videooptIMAGE—
model_patchoptMODEL_PATCH—
pose_strengthoptFLOAT0.450–2Lower this if pose skeleton lines show in the output.
pose_end_percentoptFLOAT0.550–1Stop pose control partway through sampling to let the scene clean up.

Outputs (2)

NameTypeDescription
multicam_framesIMAGE—
original_audioAUDIO—