Clean relight · multicam V4
Turn the Bedroom Take Into a Studio Session
- source_frames
- source_audio
- h3_frames
- wan_model
- wan_clip
- wan_vae
- wan_vision
- klein_model
- klein_clip
- klein_vae
- multicam_frames
- original_audio
Clean relight · multicam V4 is the pack's finishing pass: it takes the H3 multicam result you just rendered and rebuilds it in a different environment, with new lighting, while keeping the performance - timing, mouth movements, gestures - from the source.
The node's own description is unusually specific about what it wants: 24 fps source frames, the same source audio, and Wan and Klein MODEL inputs that already include their tested LoRAs. Read that twice. This node does not load models for you; it consumes models you've already prepared, and the difference between it working and producing mush is whether your Klein and Wan stacks are the ones it was tuned against.
What you connect
Required: source_frames, source_audio, h3_frames (the assembled V3 result - must be exactly the source frame count), shot_plan, look_enabled (default off), look_prompt (a grey-studio description by default), width 832 / height 480, seed, klein_seed, wan_steps 4 and klein_steps 4.
Optional, all lazy: wan_model, wan_clip, wan_vae, wan_vision, klein_model, klein_clip, klein_vae; plus the tuning knobs - prompt_preset, wan_cfg (3.0), mouth_strength (1.5), control_strength (1.0), klein_width 1024 / klein_height 576, background_colour, remove_from_source, segmentation_model, max_take_frames (120).
Outputs are multicam_frames and original_audio, passed through to the end of the chain.
Three of those knobs deserve a sentence each. remove_from_source is free text naming the old scene's objects - the tooltip suggests bedroom, bed for the reference clip - and it becomes a negative clause in both prompts. max_take_frames caps how long a single generated take may be; longer sources get split into bounded takes so nobody has to raise a GPU limit to process the whole clip. And prompt_preset switches between your wording and the Approved grey studio (reference clip) wording, which is the exact prompt from the author's reference probe, hard-coded, supporting angles 0 and 1 only.
With Look off, it costs nothing
This is the nicest design decision in the node. With look_enabled off it returns the H3 frames and audio trimmed to length, and its lazy check_lazy_status never asks for a single model input. Nothing loads. The README says it plainly: when Look is off, it returns the H3 result without loading the image or video models. So you can leave the whole relight branch wired in your graph and toggle it.
What happens when it's on
Per shot, three stages. Klein relights the shot's first frame: VAE-encode, reference-latent, CFG guider at cfg 1.0, an euler sampler on a Flux2Scheduler at your Klein resolution and step count, decode. That's your anchor - the same person in the new room.
Then a CLIP-vision encode of that anchor goes into a Wan image-to-video conditioning, and a VACE conditioning on top of it carries the control_video from Zura · Person Motion Control - the performer on flat grey, which is what stops the old room bleeding back in. control_strength sets how hard that control bites. Sampling is er_sde, simple scheduler, wan_steps steps, wan_cfg 3.
Between the two, generated angles get a mouth-timing fix: the guide frames are run through Match Mouth Motion at mouth_strength, with the source frames as the reference, since a re-angled guide's mouth can lead or lag the real speech and the video model amplifies that. Set mouth_strength to 0 and it's skipped entirely.
Shot lengths are padded to the 4k+1 grid Wan wants. Shots using the original camera and no shot prompt are generated as one continuous take, chunked by max_take_frames, then sliced by shot boundaries; everything else is a per-shot take.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/ZURAVFX/ComfyUI_zura_nodes
pip install -r ComfyUI_zura_nodes/requirements.txt
Restart. You need ffmpeg/ffprobe on PATH, the person segmentation weight at models/ultralytics/segm/person_yolov8m-seg.pt for the foreground control, and face-alignment==1.4.1 (in the requirements file) for the mouth fix. Plus, separately, a FLUX.2 Klein stack and a Wan 2.2 IDV2V stack with their LoRAs - the pack ships no weights.
Where people get burned
The shot plan must match the source frame count, and the plan's signature is checked if it has one - re-trim and you'll be told to re-plan. Every shot has to be contiguous: shots that don't start exactly where the previous one ended are rejected outright, no gaps, no overlaps, 60 shots maximum. A shot longer than max_take_frames raises and tells you to shorten it. width/height and klein_width/klein_height must be multiples of 16 between 128 and 2048. And h3_frames must be exactly the source length - connect the assembler's output, not a trimmed copy.
Finally, the caveat the README repeats and the source file states up front: structural validation is not visual sign-off. This node can assemble a technically perfect clip with the wrong face in the wrong room. Render short first.
Inputs (29)
| Name | Type | Default | Description |
|---|---|---|---|
| source_frames | IMAGE | — | |
| source_audio | AUDIO | — | |
| h3_frames | IMAGE | — | |
| shot_plan | STRING | — | |
| look_enabled | BOOLEAN | false | — |
| look_prompt | STRING | A clean neutral-grey photography studio with a seamless cyclorama and floor, soft large key light from camera left and subtle rim light. | — |
| width | INT | 832128–2048 | — |
| height | INT | 480128–2048 | — |
| seed | INT | 3141590–18446744073709550000 | — |
| klein_seed | INT | 338742110–18446744073709550000 | — |
| wan_steps | INT | 41–60 | — |
| klein_steps | INT | 41–60 | — |
| wan_modelopt | MODEL | — | |
| wan_clipopt | CLIP | — | |
| wan_vaeopt | VAE | — | |
| wan_visionopt | CLIP_VISION | — | |
| klein_modelopt | MODEL | — | |
| klein_clipopt | CLIP | — | |
| klein_vaeopt | VAE | — | |
| prompt_presetopt | COMBO | Custom look | Custom look uses your prompt for any performer. Approved grey studio restores the exact test-clip wording, including man/sweatshirt and wide-left camera; it ignores Look prompt. |
| wan_cfgopt | FLOAT | 3.001–10 | — |
| mouth_strengthopt | FLOAT | 1.500–2.5 | — |
| control_strengthopt | FLOAT | 1.000–2 | — |
| klein_widthopt | INT | 1024128–2048 | — |
| klein_heightopt | INT | 576128–2048 | — |
| background_colouropt | STRING | #929398 | — |
| remove_from_sourceopt | STRING | Old-scene objects to exclude from a new location. For this reference clip: bedroom, bed. Change it for other footage. | |
| segmentation_modelopt | STRING | segm\person_yolov8m-seg.pt | — |
| max_take_framesopt | INT | 12012–240 | — |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| multicam_frames | IMAGE | — |
| original_audio | AUDIO | — |