CV Background Subtract (Batch)
Pull a subject out of static-camera footage without a segmentation model
- images
- foreground
- background
- foreground_fraction
If the camera is bolted down, you don't need SAM, BiRefNet, or anything with weights. You need to know which pixels changed. This node takes a whole clip, learns one background, and hands you a foreground mask per frame - the honest, unglamorous approach to pulling a moving subject out of fixed-camera footage.
The vid2vid crowd has been doing this forever, and there's a reason: masking the subject separately from the environment is how you stop the model from chewing the background into mush across a clip. The old SD-CN-Animation ghosting threads land on the same advice every time - separate the moving object first, then animate it.
How it works, and why it's one node
The subtractor is created, run over the entire sequence and dropped inside this node. That's not laziness. A stateful cv2 background model can't travel through a cached graph - ComfyUI re-executes nodes based on whether their inputs changed, so the object would be updated an unpredictable number of times. Folding the clip into one execution is the clean answer, and it's the opposite choice from CV Background Model (Update), which threads the state through your wires one frame at a time. Batch when you have the whole clip handy; per-frame when you're streaming frames through a loop.
images wants a batch in frame order, from a fixed camera, at least 2 frames. More is better - the model needs to see the scene.
Picking an algorithm (the author measured this)
Seven options, and the differences are bigger than the marketing suggests. Measured on a 200-frame fixed-camera clip against a temporal-median reference that calls 5.0% of the frame moving:
- MOG2 reads 4.6%. Just use it. It's the default, it converges in a few dozen frames, it's fast, and it gives you a background plate.
- KNN is the other core OpenCV option, similar territory.
- CNT, GSOC and LSBP read 7.9 / 9.3 / 13.3% on the long clip - fine, not better. Cut the same clip to 50 frames and they jump to 16.0 / 12.5 / 16.0% while MOG2 barely moves. They aren't less accurate; they're unconverged. Those are long-sequence algorithms.
MOG is the classic mixture model and the most conservative; GMG needs lots of initialization frames and provides no plate at all (the node falls back to a temporal median so you still get background).
Inputs worth touching
warmup_frames- frames spent training the model before masks start being kept. The output has that many fewer frames. ~12 is plenty for MOG2/KNN. Don't try to rescue a short clip by raising it; the algorithm needs sequence, and spending frames on warmup just leaves fewer to keep.learning_rate- default -1 lets the algorithm pick. Set it to 0 when a subject stops moving mid-clip, or the model absorbs it and your mask fades out.detect_shadows/shadows_as_foreground- MOG2/KNN mark shadows as grey 127 instead of 255. Off (the default) keeps the mask to the object, which is what a matte wants.
Outputs
foreground is a real MASK batch, one 0/1 mask per frame - morphology and CV Keep Largest Component are the usual next steps. background is the learned plate, always BGR uint8. And foreground_fraction is the sanity check I'd wire to a text preview on every run: on static-camera footage a good result is roughly 0.05–0.2. Above ~0.3 the model hasn't settled - use a longer clip or MOG2, not a bigger warmup. Near zero means nothing moved, and your mask is empty for a reason.
Install
Manager → ComfyUI CV, or:
cd ComfyUI/custom_nodes && git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"
Python ≥ 3.12, recent ComfyUI on the V3 API. Note that several of the algorithms here (bgsegm, plus the shadow handling) live in the contrib build - keep the contrib wheel winning the cv2 import or you'll find nodes and options missing.
Where people get burned
Motion isn't the only thing that changes in a shot. Auto-exposure, a passing cloud, leaves moving, a flickering light - all of that is "foreground" to a per-pixel model, and no threshold fixes a global exposure step. Stabilise your source, or freeze the model with learning_rate 0 against a plate and accept that lighting changes will show up in the mask.
Inputs (6)
| Name | Type | Default | Description |
|---|---|---|---|
| images | IMAGE | Batch of frames IN ORDER from a fixed camera. At least 2; the more the model sees, the cleaner the background estimate. | |
| algorithm | COMBO | MOG2 (core) | MOG2/KNN (core OpenCV) are fast, adaptive, give a background plate, and converge in a few dozen frames - the right default for clip-length footage. MOG is the classic mixture model and is the most conservative. CNT is extremely fast. GSOC and LSBP are the most accurate ON LONG SEQUENCES; MEASURED, all three of CNT/GSOC/LSBP read 3-4x the reference on a 50-frame clip and only settle after several hundred frames, so prefer MOG2 unless the clip is long. GMG needs many initialization frames and provides no plate. |
| warmup_frames | INT | 00–1000 | Frames used ONLY to train the model before masks start being collected. The output then has this many fewer frames. 0 keeps every frame, including the noisy first ones. ~12 is plenty for MOG2/KNN. It cannot rescue CNT/GSOC/LSBP on a short clip: those need hundreds of frames of actual sequence, and spending them on warmup just leaves fewer to keep - check 'foreground_fraction' rather than assuming. |
| learning_rateopt | FLOAT | -1.00-1–1 | How fast the background adapts, per frame. -1 (the default) lets the algorithm choose; 0 freezes the model after the warmup, which is what you want when a subject stops moving and would otherwise be absorbed into the background; 1 makes every frame the new background. |
| detect_shadowsopt | BOOLEAN | true | MOG2/KNN only: mark shadows separately (they come out as grey 127 rather than white 255) - see 'shadows_as_foreground'. |
| shadows_as_foregroundopt | BOOLEAN | false | Count detected shadows as part of the subject. Off (the default) keeps the mask to the object itself, which is usually what a matte needs. |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| foreground | MASK | One 0/1 MASK per frame (a MASK batch): the moving / new pixels. Clean it up with morphology and 'CV Keep Largest Component'. |
| background | NPARRAY | The learned background plate, always BGR uint8. MOG2/KNN/GSOC/LSBP provide one directly; CNT's is single-channel and is promoted to BGR; MOG and GMG provide none, so the temporal median of the clip is returned instead. Preview it with 'Preview CV Array'. |
| foreground_fraction | FLOAT | Average fraction of pixels flagged as foreground (0-1) across the kept frames - the number to sanity-check every run against. On static-camera footage a good result is a small minority, roughly 0.05-0.2. Above ~0.3 the model has not settled: use a longer clip or switch to MOG2 rather than raising warmup_frames. Near 0 means nothing moved. |