CV Array Statistic
How to find the motion that isn't the moving thing
- nparray
- value
- text
Reducing a whole array to one number per channel is not a glamorous operation, and it's the hinge that a surprising number of computer-vision pipelines turn on. The example the author puts front and centre is the good one: you've computed an optical-flow field, and you want to know what the background is doing - the camera motion, the pan, the general drift. Most of the field moves that way. A moving object contaminates a minority of the pixels. Take the median of the field and you have the background's motion, robustly, without ever telling the node what's foreground.
Why you'd reach for it
- Background subtraction for motion segmentation. Get the median flow vector, subtract it from the field, threshold the residual - what's left is whatever moves differently from the scene. No training, no model, milliseconds.
- Estimating a global offset between frames. Same idea at the extreme: the median of a difference array is the systematic shift.
- Reading a measurement you already computed.
CV Create CCM Model,CV Match Image Moments,CV Array Shape, any of the pack's analysis nodes emit arrays; this turns one into a number.
It's the analysis half of a pack that keeps data and drawing rigorously separate: this node never renders anything. That's what makes it composable.
How it works
The array is reduced over every pixel, per channel, with NaNs and infinities ignored - which matters more than it sounds, since optical-flow fields and radiance maps are both happy to produce them. A 2-D [H,W] array gives you one number; a 3-D [H,W,C] array gives you C numbers, one per channel; the flow field case, [H,W,2], gives you two.
The statistic is a dropdown, and the labels carry the guidance: median is the robust dominant value, mean is the average and is dragged toward outliers, std is the spread. Also on offer: min and max. Pick by asking what a contaminating minority of pixels does to the answer - the mean moves, the median doesn't.
The inputs and outputs
nparray- a 2-D[H,W]or 3-D[H,W,C]array.NPARRAYonly; this isn't an image node.statistic- the reduction, median by default.
Outputs:
value- 1-D float64, one element per input channel. That shape is deliberate: it's the same shapeCV Scalaremits, so it broadcasts straight into the low-levelcv2arithmetic wrappers.cv2.subtract(field, median_value)subtracts the per-channel median from every pixel, which is exactly the "strip the background motion" step.text- the same numbers as a readable(a, b, ...)string. Wire it into core Preview as Text and you have a readout without thinking about it.
Install
ComfyUI Manager → ComfyUI CV, or:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
Restart ComfyUI afterwards. Requirements are the pack's usual: opencv-contrib-python-headless~=5.0.0.93, Python 3.12+, and a ComfyUI on the V3 node API. This node itself is core cv2 plus numpy, but a lot of what you'd feed it comes from contrib-backed nodes, so install the contrib wheel and don't think about it again.
Traps
- Mean vs median is not a style choice. With a large moving object - or a hand-held pan where most of the frame is moving together - the median gives you the dominant motion and the mean gives you a compromise nobody is actually moving at. Median is the default for a reason.
- Per-channel, not per-pixel-pair. For a 2-D flow field you get two numbers independently: the median of the x-components and the median of the y-components. That's almost always what you want, but it's not the median vector, and the two differ when the field has a strong rotational component.
- Empty or all-NaN input has no meaningful answer - the node ignores non-finite values, so if everything is NaN you get NaN out rather than an exception. Check
textbefore trustingvaluedownstream. - A statistic is not a registration. Subtracting the median flow stabilises the background, but it's a single global vector, not a per-pixel warp. For real alignment, that's the homography / ECC route, and in this pack that's
CV Find Homographyor the ECC exercise.
Inputs (2)
| Name | Type | Default | Description |
|---|---|---|---|
| nparray | NPARRAY | 2-D [H,W] or 3-D [H,W,C] array to reduce (e.g. an optical-flow field HxWx2). | |
| statistic | COMBO | median (robust dominant value) | Reduction applied over every pixel, per channel. Median = the dominant value (robust to outliers); mean = the average (sensitive to them); std = the spread. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| value | NPARRAY | 1-D float64, one element per input channel - broadcastable into cv2.subtract / compare / add like 'CV Scalar'. |
| text | STRING | The value(s) as a readable '(a, b, ...)' string - wire into core 'Preview as Text' (PreviewAny) to display. |