CV Array → Mask
Turning raw array data into something ComfyUI can mask with
- nparray
- mask
What it's for
ComfyUI's MASK is a very particular thing: a batch of float tensors in 0..1, one channel. Your array almost certainly isn't. It might be a uint8 binary out of a threshold, a float probability map out of a segmenter, a uint16 depth map, a bool array, or a 0..255 grayscale result from any of the pack's several hundred cv2_* wrappers.
This is the conversion that lands it in the right shape and the right range. If you've done the OpenCV work and now you need an inpaint, a mask composite, or anything else that takes a MASK socket, this is the last node in the chain. (CV Mask → CV is the reverse door, if you're heading the other way.)
How it works
One input: nparray, a single-channel [H,W] array - or a multi-channel one, which gets converted to grayscale first. Any dtype. One output: mask.
The dtype handling is the interesting part, and it's more thoughtful than a cast:
- uint8 passes straight through and gets divided by 255, so a 0/255 binary mask arrives as 0/1, which is exactly what the mask convention wants.
- uint16 is divided by 257, landing back on the 0..255 scale before that.
- bool becomes 0 or 255.
- Float arrays are rescaled from their own range into 0..255 before the divide. That's what lets a distance transform, a disparity map or a classifier's score map become a visible, usable mask rather than a black square - the values were never in 0..1 to begin with.
- Non-finite values are clamped to the finite range rather than left to poison it. A stray
nanorinf- which distance transforms and optical flow produce routinely - would otherwise make the min/max computation useless and collapse the whole mask to near-black.
A subtlety worth knowing, because it's the difference between "works" and "why is my mask soft": this is a scale, not a threshold. Anything you feed in comes out as a continuous field in proportion to its own values. If you wanted a hard, binary edge, threshold it first (cv2_threshold or one of the pack's threshold nodes) and convert the result.
Where it fits
The productive pattern is: cv2 operation → arithmetic to get the field you want → convert → mask socket. Example: run CV Connected Components or CV Find Contours, build a filled region as a mask array, convert, and feed a per-region inpaint. Another: cv2_distanceTransform → convert → use the soft field as a feathered blend weight, which is a genuinely nice trick and free of any model.
The pack's masking corner has a whole family in both directions - CV Masks to BBoxes, CV BBoxes to Masks, CV Overlay Masks, CV Labels to Masks, CV Keep Largest Component - so this node rarely travels alone.
Install
ComfyUI Manager, search ComfyUI CV. Or:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
Restart afterwards. Python ≥ 3.12, a ComfyUI recent enough for the V3 node API, and:
pip install "opencv-contrib-python-headless~=5.0.0.93"
Where people get burned
Min-max scaling makes frames inconsistent. Because float input is scaled from its own per-array range, array A with values 0..0.3 and array B with 0..0.9 both come out spanning the full mask range, at different scales. If you're processing a sequence, normalise explicitly before this node so every frame is on the same scale.
Your multi-channel array went through a BGR assumption. Three or more channels get converted to grayscale as though they were BGR, then reduced. Feeding a heatmap, a LAB array or a three-channel feature map in gives you a mask derived from a channel-weighting you didn't choose. Convert to the single channel you actually mean, upstream, with cv2_extractChannel or a raw cvtColor.
Off-by-255 thinking. A value of 255 in your array is full mask; 255 is not "a large float". If you've been doing arithmetic that produced values up in the thousands, they'll all clamp to full mask and you'll get a solid rectangle. Check with Inspect CV Data when a mask comes out as one flat tone.
It takes an array, not a LATENT. Latents aren't images; if you want a mask out of latent-space data, unwrap it first and think about whether the cell-versus-pixel units still mean what you think they mean.
One array per call. It's a single-frame converter, not a batch one. Looping over a batch is a graph-design question, and the pack's answer to that is usually the list-versus-batch policy idiom rather than running this node a hundred times.
Inputs (1)
| Name | Type | Default | Description |
|---|---|---|---|
| nparray | NPARRAY | Single-channel [H,W] ndarray (multi-channel is converted to grayscale first). Any dtype; scaled to a 0-1 mask. |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| mask | MASK | — |