Zura · Person Motion Control
Grey Out the Room, Keep the Person
- images
- control_frames
- person_masks
Zura · Person Motion Control takes a video and hands you back the performer pasted onto a flat grey background, plus the mask it used. That's it. It generates nothing, and it is not a compositing node - the description in the source is blunt about it: the result is conditioning, not a finished masked composite.
So what's it for? Changing the room. When you re-angle or relight a performer with an ID-V2V / Wan VACE style route, the model is being asked to invent a new environment. If the old environment is still visible in the conditioning, it comes back. Painting the room grey and feeding the model a clean foreground plate tells it "this is the person; the room is disposable."
Inputs
images- the frames.segmentation_model- a dropdown built at node-load time from the files in ComfyUI'sultralyticsmodel folders, filtered to names containingseg. Which is why the folder has to exist and contain something before the node is any use; it also registers themodels/ultralyticsroot and any sibling of a checkpoints folder, soextra_model_pathsinstalls work.background_colour- a six-digit hex string, default#929398(that warm grey you'll see in the pack's example prompts). A bad value raisesBackground colour must be a six-digit hex value.mask_erode_px- default 2, shrinks the mask inward. Worth a couple of pixels so you don't hand the model a fringed halo of the old room around hair and shoulders.edge_softness_px- default 1.0, a Gaussian blur on the mask, which is what stops the blend looking like a paper cutout.
Outputs: control_frames (the grey-background plate you feed a VACE or video-to-video conditioning input) and person_masks (a real MASK, useful for previewing or for your own downstream work).
The mechanism, honestly
It runs Ultralytics YOLO per frame - person class only, imgsz=960, conf=0.25 - keeps the largest mask, binarises it, then throws away any stray components with a connected-components pass so a second person in the background can't leak into your foreground. Then erode, blur, composite: frame * mask + grey * (1 - mask).
That per-frame detection is where your wall-clock goes, and where it can fail: if no person is found in a frame it raises Person not detected in frame N; cannot safely preserve performance rather than silently producing a grey frame. Heavy occlusion, a performer half out of frame or an extreme angle will do it.
Two things to know before you build on it. It is deliberately pre-generation: it strips old-room information from the conditioning only, and the frames you finally deliver are always the full VAE decode. And the detector runs through ultralytics - the same AGPL-3.0 territory that covers its weights, the same dependency that reached ComfyUI users in the 2024 poisoned-release incident. Fine for personal work; think about it before you ship a product.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/ZURAVFX/ComfyUI_zura_nodes
pip install -r ComfyUI_zura_nodes/requirements.txt # brings ultralytics, face-alignment, requests, yt-dlp
Then put a person segmentation weight where ComfyUI looks for ultralytics models - models/ultralytics/segm/person_yolov8m-seg.pt is the name the pack's other nodes expect, and Impact Subpack can fetch it for you. Restart ComfyUI after dropping the file in: the dropdown is populated when the node's inputs are built, so a model added later won't appear until you reload.
In the V4 clean relight node this runs internally, with mask_erode_px pinned to 2 and edge_softness_px to 1.0 - the same values you'd probably pick by hand.
Inputs (5)
| Name | Type | Default | Description |
|---|---|---|---|
| images | IMAGE | — | |
| segmentation_model | COMBO | 0 options: | |
| background_colour | STRING | #929398 | — |
| mask_erode_px | INT | 20–20 | — |
| edge_softness_px | FLOAT | 1.000–10 | — |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| control_frames | IMAGE | — |
| person_masks | MASK | — |