Nodes/Zura Nodes/Zura · Match Mouth Motion
ComfyUI Node

Zura · Match Mouth Motion

Fix the Guide's Mouth Before the Model Ever Sees It

By ZURAVFX·Created 22 days ago·Updated about 16 hours ago· 1
Zura · Match Mouth Motion
  • source_frames
  • angle_guide_frames
  • mouth_matched_guide
  • mean_aperture_error_before
◄source_start_frame0►
◄strength1.5►
◄enabledtrue►

Zura · Match Mouth Motion is the most specific node in this pack and, if you're doing the ID-V2V relight route, one of the more satisfying. It takes your original footage and the camera-angle guide you generated from it, finds the faces in both, measures how open the mouth is in each, and warps only the guide's mouth region until the apertures line up.

Then the video model gets a guide whose lips already follow the source speech.

This is a pre-generation fix, and that distinction is the whole idea. The node's own docstring is explicit that source pixels are never composited into the result - you're not pasting a mouth, you're reshaping the guide's own mouth so the model has less to get wrong. Fixing lip sync after the fact means another pass over every frame; nudging the guide costs a few seconds per frame and happens once.

Inputs and outputs

Required: source_frames (the original 24 fps clip), angle_guide_frames (the generated camera guide), source_start_frame (default 0 - where in the source the guide starts), strength (default 1.5, range 0–2.5, step 0.1) and enabled (default true).

Outputs: mouth_matched_guide (the reshaped guide image sequence) and mean_aperture_error_before - a float. That second output is the useful one no one expects.

What it actually does

Face-alignment's 2D landmark detector finds 68 points in the source frame and in the guide frame. From those it computes an aperture ratio: inner lip gap divided by mouth-corner width. The difference between source and guide is the error.

Then it warps the guide. Not a paste - a smooth displacement field centred on the mouth, with the envelope falling off steeply horizontally (exp(-0.5*dx⁴ ...)) and a tanh term that pushes the upper and lower lips in opposite directions, so the mouth opens or closes rather than smearing. The applied change is strength × (desired − current) / 2 per lip. cv2.remap with cubic interpolation and border replication does the resampling. Only the mouth moves; the rest of the frame is untouched.

mean_aperture_error_before is the mean absolute aperture difference across the clip before the correction. If it comes back near zero, the guide already matched and strength is doing nothing. If it's large, the guide's mouth is badly out of step and you should expect the render to have leaned on the model's imagination. It's a real diagnostic, and it's the only one you get - there's no "after" number, which is a slightly odd asymmetry but the useful half is the one you have.

Install, and the part that will annoy you

cd ComfyUI/custom_nodes
git clone https://github.com/ZURAVFX/ComfyUI_zura_nodes
pip install -r ComfyUI_zura_nodes/requirements.txt

That pulls face-alignment==1.4.1 - pinned, not a range, and it's the dependency people trip over, because face-alignment drags in old torchvision/numba plumbing and behaves badly on some environments. The landmark weights download on first use into the node's own model_cache/hub folder, so the first run is slower and needs internet.

Then the surprise that is actually a decision: it runs on CPU, deliberately. The author's comment in the source says GPU FAN landmarks vary by a few mouth pixels between fresh runs and the downstream ID-V2V amplifies that difference - CPU was byte-repeatable. So this node costs you wall-clock, on purpose, for reproducibility. It also caps PyTorch to four threads while running; the comment reports a 21-frame probe at roughly 39 seconds with four threads versus about 90 with more.

Errors you'll meet

Mouth guide needs N source frames starting at frame X; only M are available means your source_start_frame and guide length don't fit inside the source clip - usually because you're feeding a trimmed shot from a longer clip. Mouth landmark detection failed at frame N means no face was found in either the source or the guide at that frame; profile angles and heavy motion blur are the usual causes. With enabled off, or strength at 0, it returns the guide untouched and an error of 0.0, which is a cheap way to A/B the whole node.

CategoryZura/Video Control

Inputs (5)

NameTypeDefaultDescription
source_framesIMAGE—
angle_guide_framesIMAGE—
source_start_frameINT00–100000—
strengthFLOAT1.50–2.5—
enabledBOOLEANtrue—

Outputs (2)

NameTypeDescription
mouth_matched_guideIMAGE—
mean_aperture_error_beforeFLOAT—