Nodes/comfyui-svdint4/Video Latent Composite Masked
ComfyUI Node

Video Latent Composite Masked

Put the upscaled pass back without losing the mask

By wjie98·Created 3 months ago·Updated 3 days ago· 3
Video Latent Composite Masked
  • original_latent
  • replacement_latent
  • mask
  • latent
◄typeminimax►

Second-pass video work is where the plumbing gets ugly. You sampled once, you like 80% of it, you upscaled the latent, you'd like to re-diffuse the part you didn't like at the higher resolution - and now you're holding two [B,C,T,H,W] tensors that need to be welded into one, with a mask that survives the weld so the next sampler only touches what you asked it to.

That's this node. It composites two matching video latents under hard coverage and hands you back a latent whose noise_mask describes exactly the region it replaced. It's the counterpart to the pack's Set Video Latent Noise Mask: one sets a mask, this one applies a composite and preserves it.

How the coverage is decided

original_latent is the clean original - your high-resolution source, whose samples are kept wherever the mask is zero. replacement_latent is the thing you want to graft in: typically the upscaled first pass, i.e. a denoised_output. Its own inherited noise_mask is the starting point.

Then, if you also wire mask, the extra [frames,H,W] or [B,frames,H,W] mask is mapped into latent time and unioned with the inherited coverage. The type combo picks the temporal mapping, and the choices are the same profiles Set Video Latent Noise Mask offers: wan, minimax (the default, for H3 video), ltxv, hunyuan_video, hunyuan_video_15, and mochi. Each VAE compresses time differently, so a mask authored per-pixel-frame has to be merged differently - H3, for example, folds [1,4,4,4,4]*n+[1,4] groups into latent positions, which is why 124 frame masks with the first five black become 37 latent masks with the first two black. The tooltip is precise about scope: type is used only to map that additional image-frame mask. An inherited latent mask is never retimed.

The output is one latent. samples become replacement where coverage is set and original everywhere else - exactly, no blending. noise_mask becomes that same binary coverage, expanded to [B,1,T,H,W]. Metadata comes from original_latent, but its mask is gone, replaced by the composite's.

The rules that will actually stop you

Every positive value means full replacement. Not opacity, not strength - coverage. Threshold or harden a soft/noisy mask upstream if that's not what you want, because a fuzzy 0.3 mask becomes a hard edge here.

Shapes must match exactly. Same [B,C,T,H,W], channel for channel. There's no latent resizing, no retiming, no feathering; a mismatch is an error telling you to align the two first. If your two latents came out of different crops or different frame windows, fix that before it reaches this node.

Standalone latents only. H3's native AV latent is a nested object, not a [B,C,T,H,W] tensor, so split it with H3 Separate AV Latent first and rejoin with H3 Concat AV Latent afterwards. Same story for LTX-2-style separated latents.

Inherited masks have to fit. If replacement_latent carries a noise_mask, it must match its own T/H/W with batch 1 or B and channels 1 or C.

Where it fits

The intended shape is a masked second pass: composite the upscaled result onto the original, then run the continuation sampler on the clean composite with native RandomNoise and the remaining sigmas. Note the pack's own caveat - the H3 Add Noise node isn't made mask-aware by this one, so the clean-composite-plus-noise route is the supported path. The other everyday use is restoring what the upscaler didn't claim: a latent upscale enlarges an inherited mask by conservative spatial maximum coverage, so compositing back onto the original is how you keep the untouched area genuinely untouched instead of overwriting the union with a fresh mask.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/wjie98/comfyui-svdint4.git
# restart ComfyUI

Listed as comfyui-svdint4, README titled "ComfyUI Turing Utils", cloning comfyui-turing-utils - the project was renamed, it's one pack, and the folder name is irrelevant. Pure Python and torch; no models, no dependencies beyond what ComfyUI ships, and the README's CUDA kernel build is only needed for the pack's quantisation and attention nodes.

CategoryTuring Utils/video

Inputs (4)

NameTypeDefaultDescription
original_latentLATENTClean original high-resolution video [B,C,T,H,W]. Mask=0 preserves these samples. Its metadata is retained, but its noise_mask is replaced.
replacement_latentLATENTClean replacement video, e.g. upscaled first-pass denoised_output. Shape must match original_latent. Its noise_mask is inherited; mask=1 uses these samples.
typeCOMBOminimaxSame temporal mapping as Set Video Latent Noise Mask. Only used to map additional image-frame masks, never to retime inherited latent masks.
maskoptMASKAdditional [frames,H,W] or [B,frames,H,W] mask. Its nonzero coverage is unioned with the inherited mask; even small positive values mean full replacement.

Outputs (1)

NameTypeDescription
latentLATENT—