Nodes/ComfyUI-MatAnyone2/MatAnyone2 Video Matting
ComfyUI Node

MatAnyone2 Video Matting

Cut a person out of a whole video from one mask on the first frame

By starsFriday·Created 7 months ago·Updated 7 months ago· 0
MatAnyone2 Video Matting
  • matanyone2_model
  • images
  • first_mask
  • foreground
  • composite
  • alpha
◄warmup10►
◄erode_kernel10►
◄dilate_kernel10►
◄mask_threshold0.50►
◄max_internal_size-1►

You paint one rough mask over the person in the first frame, and this node returns a per-frame alpha matte for the entire clip. That's the whole pitch, and it's a big deal if you've ever tried to background-remove a video the naive way: running an image model like BiRefNet on every single frame. That works, and it flickers, and it's slow, and every frame is guessing fresh. Video matting instead propagates the mask forward through the sequence, using the temporal information an image model throws away. MatAnyone2 is the current name here - the community has been excited about the approach since the v1 release in early 2025, largely because it handles hair edges that hard-matte segmentation (SAM2 included) just blurs into mush.

How it works

The node is a thin wrapper around MatAnyone2's InferenceCore, with a memory-propagation loop. The details that matter to you:

  • warmup (default 10) is first-frame warmup iterations. Before touching your real frames, the model replays frame 0 with the mask over and over, building up its memory bank so the actual sequence starts from a confident state. Crank it up if the first seconds of your clip bleed or flicker; keep it low for long clips where speed matters.
  • Your mask is binarized at mask_threshold (default 0.5), then dilated, then eroded by dilate_kernel / erode_kernel (both default 10). That cleans up a sloppy first-frame scribble. Beware the defaults: 10 pixels of morphology is a lot, and thin strands of hair can get eaten. If hair comes out too thin or missing, drop both toward 0–3; if the matte bleeds into the background, raise them.
  • max_internal_size (default -1) caps the internal shortest side of the frames. Leave it unless you're hitting VRAM limits - then it's your cheap resolution dial: the model downscales internally and quality drops gracefully instead of OOMing.

Internally it's the single-target pipeline (objects=[1]), so it's built for "this person in the foreground," not multi-object decomposition. Run it once per subject.

Inputs and outputs

Three inputs, and two of them are the point:

  • images - an IMAGE batch [T, H, W, 3], straight from any video loader node. Frames get clamped to 0–1 and run through in order.
  • first_mask - a MASK. White = subject, black = background. The white/black convention matters, and if you feed multiple mask frames, only the first one is used.
  • matanyone2_model - the output of the MatAnyone2 Model Loader.

Three outputs, and the last one is the one you actually want:

  • foreground - the cutout with a black background. Good for a quick look.
  • composite - the same cutout composited over a green screen, matching the official inference visualization. It's a preview, not your deliverable.
  • alpha - the per-frame matte. This is the real output: wire it into any composite-over-background node to drop the person onto a new scene, or run it through a feather/blur to soften edges. The foreground output is really just frame * alpha computed for you.

Installing and getting the frames in

Same install as the loader - ComfyUI Manager (search "MatAnyone2"), or:

cd /path/to/ComfyUI/custom_nodes
git clone https://github.com/starsFriday/ComfyUI-MatAnyone2.git
cd ComfyUI-MatAnyone2
pip install -r requirements.txt

Restart, and let the loader auto-download matanyone2.pth on first run (see the loader article for the model-folder gotcha: ComfyUI/models/matanyone2/, not checkpoints).

For the first-frame mask, the standard move is to freeze a frame, run it through a segmentation node - SAM, GroundingDINO, or a BiRefNet portrait model - and feed that mask in. Don't obsess over a perfect first frame: the morphology passes plus warmup exist precisely because the model can take a rough starting mask. That said, a messy scribble will cost you.

Troubleshooting

  • Flicker in the first seconds → raise warmup; the memory bank wasn't confident yet.
  • Hair evaporating → lower erode_kernel / dilate_kernel; 10/10 is aggressive.
  • VRAM errors on big clips → set max_internal_size to something like 480 or 540 rather than dropping the whole clip.
  • Mask ignored → check your first_mask isn't inverted (white should be the subject). Wrong-resolution masks get resized internally, so shape mismatch is handled; black-on-white is not.
  • Slow first run → expected; the model load is cached after that.

One last take: if your clip is ten frames of a static subject, just remove background per-frame with BiRefNet and stop reading. Video matting earns its keep the moment there's motion, occlusion, or a camera move - that's when frame-by-frame segmentation starts hallucinating edges and this starts looking like sorcery.

CategoryMatAnyone2

Inputs (8)

NameTypeDefaultDescription
matanyone2_modelMATANYONE2_MODEL—
imagesIMAGE—
first_maskMASK—
warmupINT100–64—
erode_kernelINT100–128—
dilate_kernelINT100–128—
mask_thresholdFLOAT0.500–1—
max_internal_sizeINT-1-1–4096—

Outputs (3)

NameTypeDescription
foregroundIMAGE—
compositeIMAGE—
alphaMASK—