CV Temporal Reduce (Background Plate)
Delete Everyone From the Shot
- images
- plate
Keep the camera still, film for a minute, and take the median of every pixel across all those frames. Anything that moves - pedestrians, cars, your friend wandering through frame - occupies any given pixel only a minority of the time, so it loses the vote. What's left is the empty scene: the background plate.
This is the oldest trick in computational photography and it still beats anything learned. It's also the foundation of fixed-camera compositing, background subtraction, and the "clean plate" that every rotoscoping decision downstream depends on.
How it works
images takes a whole ComfyUI IMAGE batch (or a LATENT batch), and the reduction runs per pixel over the batch axis. Frames are stacked, reduced in float64, and for images the result is rounded and clipped back to a BGR uint8 plate. One output pixel per input pixel.
statistic picks the reduction and it changes the meaning:
- median (default) - robust static background. Movers drop out, because the median doesn't care about outliers.
- mean - smoother, and it ghosts anything that moves. Every mover gets averaged in as a translucent smear.
- min / max - the darkest or brightest value each pixel ever took. Useful for picking the least-obstructed or brightest-sample frame, and it's a genuinely handy trick for heavy traffic, but it's not a clean plate in the usual sense.
Feed it a LATENT batch instead and the same reduction happens in float32 latent space, coming back as a (H,W,C) float NPARRAY with any channel count. That's the interesting option for diffusion-side work - latent values never get uint8-quantised, so you keep the precision, and you can build a plate without a decode round-trip.
The plate is a natural feed for CV Background Model (Update) - wire it into the mean input with learning_rate at 0, and you have a static background to diff each frame against. The pack's CV Foreground Mask (Video) blueprint is the other half of that chain.
What makes it work, and what breaks it
The subjects have to move across the clip. Someone who stands still for the whole take is in every frame, so they win the vote and become part of the plate - there's no algorithm to save you from that. Same for a parked car, or a camera that drifts: a micro-jitter in the tripod smears everything and your plate goes soft.
Memory is the real constraint, and the source says why: frames are stacked and reduced in float64. That's 8 bytes per channel per pixel before anything else happens, so 100 frames of 1080p colour is already several gigabytes resident. Keep the batch in the low hundreds of frames and consider running the plate at half resolution - a slightly soft background plate is usually fine for masking and compositing.
Also make sure your frames are actually spread across the clip. Ten consecutive frames a tenth of a second apart see movers in nearly the same place; ten spread over the take see them in different places, which is the whole mechanism.
Install
ComfyUI Manager → search ComfyUI CV (publisher bmad4ever), or:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"
Restart afterwards. Python ≥ 3.12 and a recent ComfyUI built on the V3 node API. Keep the contrib wheel: a plain opencv-python installed over it shares site-packages/cv2, empties the contrib submodules, and quietly breaks other nodes in the pack. tools/repair_opencv_contrib.py --check / --apply handles it.
28_background_subtract_bgsegm.json and exercise_background_subtraction.json use this node directly; the CV Background Plate (Video) blueprint in docs/subgraphs.md wraps the same idea. Load your clip with core's Get Video Components (or a video loader) to get a batch. 01_install_example_inputs.json copies the sample media in - run it once, then reload the page.
Where it bites
A plate that comes back as a single-pixel image, or zeroed, means you passed an empty batch - the node keeps the workflow alive with a blank plate rather than erroring. Check your loader.
And the standard pack disclaimer, straight from the README: personal, heavily LLM-assisted project, updates not planned, workflows tuned to specific datasets and not production-grade. This node's core is a one-line numpy reduction and behaves exactly as advertised; the useful part of that warning is that the pipelines built around it in the repo are demos, not templates.
Inputs (2)
| Name | Type | Default | Description |
|---|---|---|---|
| images | IMAGE,LATENT | Video frames to reduce over time: an IMAGE batch (e.g. 'Get Video Components' images) or a LATENT batch ({'samples': [B,C,H,W]}). The reduction runs over the batch dimension, one output pixel per input pixel. A LATENT is reduced in float32 latent space (values untouched, any channel count); an IMAGE reduces to BGR uint8. | |
| statistic | COMBO | median (robust static background) | Per-pixel reduction over time. Median = robust static background (moving objects drop out); mean = smooth average (ghosts movers); min/max = darkest/brightest. |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| plate | NPARRAY | The reduced background plate - BGR uint8 (H,W,3) for an IMAGE batch, or FLOAT32 (H,W,C) for a LATENT batch. Wire into 'CV Background Model (Update)' as 'mean', or preview it with 'Preview CV Array'. |