Set Video Latent Noise Mask
Your mask has 124 frames and your latent has 37. This node is the conversion.
- samples
- mask
- latent
What it's for
A video latent is 5D: [B, C, T, H, W]. A mask you paint or generate is 3D: [frames, H, W], at pixel resolution. Core ComfyUI's Set Latent Noise Mask wants a mask that already agrees with the latent's own grid, including its T - so knowing the VAE's temporal compression is your job. This node takes the mask your filmstrip actually produced, folds it into latent time for the model you name, and refuses to guess when the numbers don't work out.
Reach for it when you want part of a clip held fixed and the rest redrawn: erasing an object a YOLO/SAM pass found, repainting a face across a shot, keeping a logo and background locked while the subject changes. You could describe the edit to Bernini instead - maskless video editing is the 2026 answer for a lot of this - but a mask is still the only way to keep unmasked pixels bit-identical.
How the mechanism works
noise_mask in ComfyUI is the sampler's permission slip: 1 means "this latent cell gets redrawn from noise", 0 means "leave it alone". So the only hard problem is getting a pixel-frame mask into latent time.
If your mask frame count already equals the latent's T, the node copies it straight through and ignores the type widget entirely. If it doesn't, it merges consecutive pixel frames into the latent's temporal groups for the VAE you picked:
wan,hunyuan_video,hunyuan_video_15- first frame alone, then groups of 4ltxv(LTX-Video and LTX-2 video) - groups of 8mochi- groups of 6minimax(H3) - the odd one,[1,4,4,4,4]*nfollowed by[1,4], which is why a 124-frame H3 clip ends up as 37 slices
Each group collapses with a maximum, in time and in space: if any frame in the group is white at that spot, the whole latent cell counts as redraw. Small regions survive and nothing gets averaged away, but the redraw area grows to the union of the group - no per-frame surgical inpainting, the latent has no per-frame resolution to give you. Feather your mask and the strong edge wins; the node does no blurring of its own.
The inputs and output
Three required inputs, one output, and you only really think about two of them.
- samples - a standalone 5D video latent. Images and audio latents are rejected outright.
- mask -
[frames, H, W], or[B, frames, H, W]for a different mask per video in the batch; a 3D mask is shared across the batch. Values must be finite and inside[0, 1]. It does not have to match the latent's spatial size - the node max-pools it down toH_lat, W_lat, so a full-resolution filmstrip is normal. - type - the VAE profile above, default
minimax. It only matters when the frame counts differ, which is when getting it wrong hurts.
latent is the only output, and it goes into KSampler.latent_image. Your samples tensor and metadata pass through untouched; any existing noise_mask is replaced.
Installing
It ships inside the Turing Utils pack. ComfyUI Manager, search the pack title (comfyui-svdint4, sometimes listed as Turing Utils), or:
cd ComfyUI/custom_nodes
git clone https://github.com/wjie98/comfyui-svdint4
Restart ComfyUI. No model downloads and no CUDA kernel build are needed for this node. The README spends most of its length on a separately installed comfyui-turing-utils-kernel pip package for the pack's ConvRot and sparse-attention nodes - separate lifecycle, and this node never touches it. Its requirements.txt deliberately contains only safetensors so Manager won't start compiling CUDA at you. Keep ComfyUI reasonably current, though - the node is written against ComfyUI's V3 node schema (comfy_api.latest), so a stale install fails to register the pack.
Where people get burned
The frame-count error is the whole experience. Expected 37 latent-frame masks or 124 image-frame masks. No temporal interpolation, padding, or truncation is performed. - the count matches neither T nor the profile's pixel-frame grid for your clip length. The classic 81-frame Wan clip is T=21, so it wants 21 masks or 81, nothing between. And because type is only consulted when counts differ, a wrong profile maps masks to the wrong slices silently: 17 masks against a T=3 latent is valid for ltxv and an error for wan.
Packed audio/video latents are not accepted. On LTX-2 or H3 you get a nested AV latent, and the node tells you to split it first (LTXV Separate AV Latent, or the pack's own H3 nodes) and concatenate after masking.
Mask batch mismatch. A [B, ...] mask batch must be 1 or exactly the video batch size. No broadcasting from 3 to 2.
Masks from math chains. Values slightly outside [0, 1], or a NaN inherited from an upstream composite, are hard errors, not clamped. Fix it upstream.
Last thing worth knowing: this node maps temporal groups, not the VAE's full receptive field, so pixel-identical decoded boundaries aren't guaranteed, and it doesn't make H3's Add Noise node mask-aware.
Inputs (3)
| Name | Type | Default | Description |
|---|---|---|---|
| samples | LATENT | Standalone video latent [B,C,T,H,W]. Separate audio/video latents first. | |
| mask | MASK | [frames,H,W], or [B,frames,H,W] for separate video batches. A single mask sequence is shared across the video batch. 0 preserves, 1 redraws. | |
| type | COMBO | minimax | Used only when frame counts differ. wan (2.1/2.2), hunyuan_video and hunyuan_video_15: first frame then groups of 4; ltxv (LTX-Video/LTX-2 video): groups of 8; mochi: groups of 6. minimax (H3 video): [1,4,4,4,4]*n+[1,4]. |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| latent | LATENT | — |