Nodes/ComfyUI CV/CV Background Model (Update)
ComfyUI Node

CV Background Model (Update)

One frame, one node, no hidden state

By bmad4ever·Created 3 months ago·Updated 14 days ago· 1
CV Background Model (Update)
  • frame
  • mean
  • variance
  • mean
  • variance
  • foreground
◄learning_rate0.05►
◄threshold2.5►
◄init_variance100.00►
◄min_variance16.00►

Static-camera footage has a cheat code, and it isn't a segmentation model. If the camera doesn't move, "what changed since the last frame" finds your subject for roughly zero compute - no weights, no ONNX, no license questions. That's what this node is: one step of temporal background subtraction, where the model itself travels down your wires.

The reason it's shaped like this is worth understanding, because it explains half the design decisions in this pack. OpenCV's background subtractors are stateful objects: you call apply() frame after frame and the object remembers the scene. ComfyUI caches node outputs and re-runs only what changed, so a stateful object hidden inside a node would silently double-update, or update out of order, or not at all. So the state is the data: mean and variance go out one side and come back in the next frame's node. Thread them node to node the way you'd thread a trajectory.

How it works

It's a single-Gaussian-per-pixel model - MOG2 minus the mixture. Each pixel has a running mean and variance; a pixel is flagged as foreground when its distance from the mean, averaged over colour channels and expressed in standard deviations, beats threshold. Then the model eases toward the current frame at learning_rate.

Two inputs carry the state: mean and variance, both optional. On the first frame, leave them unwired - the model seeds from the frame, and foreground comes out empty because there's no history to compare against yet. That's deliberate, and it means you don't need a success gate to avoid a garbage first frame; there just isn't one.

The knobs

  • learning_rate - how fast the background follows the scene. 0 freezes it after the first frame, ~0.01–0.1 adapts to lighting while remembering a moving object for many frames, 1 makes every frame the new background (useless).
  • threshold - the foreground cutoff in standard deviations. ~2–3 is the working range; raise to suppress noise, lower to catch faint motion. Because it's in std devs it's scale-free.
  • init_variance / min_variance - how tolerant the model starts, and the floor that stops a perfectly still pixel flagging on 1–2 grey levels of codec jitter.

foreground is a uint8 0/255 mask, one channel. From there: threshold it, clean it up with morphology, or turn it into a MASK for the pack's overlay / array-to-mask nodes.

Two ways to run it

Frame-by-frame with the state threaded through the graph - flexible, works with per-frame nodes, but the loop has to carry two accumulators.

Or skip the chain for a fixed camera: precompute one background plate with CV Temporal Reduce, wire it into mean, set learning_rate 0, and each frame is scored independently against that plate. Now the loop carries only the output frames, which is what lets the whole detector run inside an Inspire foreach. This is the version I'd reach for on a locked-off shot.

The latent-space trick, and its trap

frame, mean and variance accept a LATENT and echo that format, so you can run the detector on VAE-encoded frames and never decode just to find motion. The trap is that the variance defaults are in uint8-intensity-squared units, and latents live at roughly unit scale. With the defaults nothing ever flags - the model is so tolerant that every pixel is background. For a latent frame, drop init_variance to about 1 and min_variance to about 0.01. threshold needs no change; it's in std devs.

Install

Manager → ComfyUI CV, or:

cd ComfyUI/custom_nodes && git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"

Python ≥ 3.12, recent ComfyUI (V3 node API). The pack pulls numpy and torch too, but you already have both.

Where people get burned

Ghosting. The model learns everywhere at learning_rate, so an object that stops moving gets absorbed into the background within a few dozen frames and stops being foreground. If a subject sits still, lower the rate. And there's no found output here to branch on - foreground is legitimately empty on frame one, so treat "empty" and "nothing moved" as the same case.

Categoryimage/CV/segmentation

Inputs (7)

NameTypeDefaultDescription
frameCOMFY_MATCHTYPE_V3Current frame: an ndarray (feed a ComfyUI IMAGE through 'Image -> CV Array', BGR uint8) or a LATENT (a single VAE-encoded frame). Compared against the running model; the model threads in whatever format arrives.
learning_rateFLOAT0.050–1How fast the model follows the scene, per frame (alpha): mean += alpha*(frame-mean). 0 freezes the model after the first frame; ~0.01-0.1 adapts to lighting while remembering a moving object for many frames; 1 makes every frame the new background.
thresholdFLOAT2.50–20Foreground cutoff in standard deviations: a pixel is foreground when its per-channel-averaged distance from the mean, divided by the model's standard deviation, exceeds this. ~2-3 is typical; raise to suppress noise, lower to catch faint motion. Scale-free (it is in std devs), so it needs no change in latent space.
init_varianceFLOAT100.000.001–100000Starting per-pixel variance (uint8 intensity^2 units; 100 = std 10) used to seed the model on the first frame and whenever 'variance' is unwired. Larger = more tolerant until the model settles. For a LATENT frame this is far too large (latents are ~unit scale): drop it to ~1 or the model starts hopelessly tolerant.
min_varianceFLOAT16.000.001–100000Floor on the variance (uint8 intensity^2 units; 16 = std 4) so a perfectly still pixel keeps a little noise budget and does not flag on 1-2 grey-level jitter (compressed / decoded video). For a LATENT frame lower it to ~0.01 - the uint8 floor of 16 (std 4) dwarfs latent deviations, so with the default NOTHING flags.
meanoptCOMFY_MATCHTYPE_V3Running per-pixel mean from the previous frame's node (same format as 'frame'). Leave UNWIRED on the first frame to seed the model from 'frame'.
varianceoptCOMFY_MATCHTYPE_V3Running per-pixel variance from the previous frame's node. Unwired -> reset to 'init_variance'.

Outputs (3)

NameTypeDescription
meanCOMFY_MATCHTYPE_V3Updated per-pixel mean (float32, same shape and format as the frame) - wire into the next frame's node.
varianceCOMFY_MATCHTYPE_V3Updated per-pixel variance (float32, same format as the frame) - wire into the next frame's node.
foregroundNPARRAYuint8 0/255 foreground mask (moving / new pixels), one channel. For a LATENT frame it is at LATENT resolution (H/8, W/8) - upscale x8 (nearest) to overlay on pixels. Empty on the first frame. Convert to a MASK for overlay / morphology.