Zura V3 · H3 multicam render
Multicam Inside MiniMax H3
- source_frames
- source_audio
- stage
- reference_video
- moge_geometry
- model
- clip
- video_vae
- audio_vae
- scene_video
- multicam_frames
- original_audio
Zura V3 · H3 multicam render is the node that makes the pack's multicam idea real: you give it a source performance and a shot plan, and it generates the clip again from each angle you picked - keeping the performer, the speech and the source audio - then hands back a single assembled sequence.
The key word in the display name is H3. This routes through MiniMax H3, the 33B omni-modal video model that generates picture and audio as one thing. That's why the lip sync has a chance here at all: the node feeds H3 the trimmed source audio as an anchor, not just the reference video, because (in the source's own comment) reference-video audio is only a reference token and H3 will otherwise freely invent a different spoken take. One line of code, and it's the difference between a re-shot and a re-voiced clip.
Worth knowing before you commit: MiniMax H3's community licence geofences the weights out of the US, EU, UK and Korea. The hosted API is global; the local weights are not, for everyone.
What you connect
Required: source_frames and source_audio (the 24 fps clip and its own audio), shot_plan (the JSON from Zura V3 · Pick camera angles), stage, seed (default 42), steps (default 8), and prompt - the default is a carefully written instruction telling H3 that <Picture 1> defines only the camera angle, to keep the exact performer, face, clothes, gestures and speaking performance from <Video 1>, and not to replace the performer.
Everything expensive is optional and lazy, which is the node's whole trick:
reference_video- your resized source frames. This is the motion reference.moge_geometry- MoGe monocular geometry for the camera warp guide. Drop it if you setuse_camera_warp_guideto false.model,clip,video_vae,audio_vae- the H3 setup, with its LoRAs already applied to the model.scene_video,look_enabled,background_mode- V3 keeps the source environment; ask for a background swap here and it raisesV3 preserves the source environment. Use V4 for a generated scene swap.render_selected_intervals_only- forces per-shot takes instead of shared full-clip takes.use_camera_warp_guide- default on.
Outputs: multicam_frames (IMAGE, the assembled clip) and original_audio (AUDIO - the source audio trimmed to exactly the frame count at 24 fps). Pipe the two into whatever save/encode node you use.
How it works
This node is a graph expander, not a sampler. It emits native nodes and returns them.
It renders only the angles your plan uses - one take per angle, not per shot, unless the shot has its own prompt or you asked for intervals only. Generated angles go through the saved angle still (from the planning pass), then a CrossViewWarp warp of the reference video to that azimuth and distance on top of the MoGe geometry, then a MiniMax H3 reference-to-video conditioning with the still as <Picture 1>, the sliced reference video as the motion reference and the trimmed audio. Then the guided latent is sampled with res_multistep, a simple scheduler and a basic guider, and decoded.
The nine positions are azimuths of −30/0/+30° across three distances, with a row-dependent vertical lens shift used purely for headroom. The code is candid that the warp's dolly distance is unreliable, so the close row doesn't zoom - it takes a native centre crop at 5/6 of the frame and scales back up with Lanczos. A repeatable crop beats asking the model to invent head pixels it never saw.
Then it wires an internal Zura V3 · Assemble shots, which slices full-clip takes into their shots and gives every original-camera shot back as untouched source frames.
The frame-count trap worth knowing
H3's valid clip lengths sit on a 17k+5 grid, and the MiniMax reference and guide nodes will silently crop an invalid length down to the previous boundary. The source comment records the actual damage: a 21-frame shot became a five-frame performance reference. The fix in the code is to pad the target latent to a valid length and repeat the last frame of each guide so all the conditioning keeps the real frames. If your short shots lose motion out of nowhere, this is why.
Install and the model files
cd ComfyUI/custom_nodes
git clone https://github.com/ZURAVFX/ComfyUI_zura_nodes
pip install -r ComfyUI_zura_nodes/requirements.txt
Then the parts that aren't in the repo: MiniMax H3 support nodes (MiniMaxH3ReferenceToVideo, MiniMaxH3AddGuide), ComfyUI-CrossViewWarp for the warp guide, KJNodes, VideoHelperSuite, and MoGe geometry if you leave the warp guide on. No weights ship with the pack - your H3 UNet, text encoder, VAEs and tested LoRAs are your own files. ffmpeg/ffprobe go on PATH.
Errors that mean something
The source video changed. Re-run angle planning. - the plan's signature no longer matches the frames, because you re-trimmed. Connect H3 model setup and warp preparation for Render multicam. - something in the required set is missing for the pass you're in. Shot N uses the original camera. Select a generated angle to use a motion prompt. - you put a per-shot prompt on a 0 shot. And The resized reference video is shorter than the shot plan. - your reference frames don't cover the whole clip.
Last thing: a valid graph proves nothing about the picture. Render thirty seconds before you render ten minutes, and watch the joins.
Inputs (18)
| Name | Type | Default | Description |
|---|---|---|---|
| source_frames | IMAGE | — | |
| source_audio | AUDIO | — | |
| shot_plan | STRING | — | |
| stage | ZURA_MULTICAM_STAGE | — | |
| seed | INT | 420–18446744073709550000 | — |
| steps | INT | 81–60 | — |
| prompt | STRING | crossview. <Picture 1> defines only the selected camera angle. Keep the exact performer, face, clothes, gestures and speaking performance from <Video 1>, with lips faithful to its original speech. Do not replace the performer. | — |
| reference_videoopt | IMAGE | — | |
| moge_geometryopt | MOGE_GEOMETRY | — | |
| modelopt | MODEL | — | |
| clipopt | CLIP | — | |
| video_vaeopt | VAE | — | |
| audio_vaeopt | VAE | — | |
| scene_videoopt | IMAGE | — | |
| look_enabledopt | BOOLEAN | false | — |
| render_selected_intervals_onlyopt | BOOLEAN | false | — |
| use_camera_warp_guideopt | BOOLEAN | true | — |
| background_modeopt | STRING | Keep source background | — |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| multicam_frames | IMAGE | — |
| original_audio | AUDIO | — |