Nodes/ComfyUI CV/CV Array → Mask
ComfyUI Node

CV Array → Mask

Turning raw array data into something ComfyUI can mask with

By bmad4ever·Created 3 months ago·Updated 14 days ago· 1
CV Array → Mask
  • nparray
  • mask

What it's for

ComfyUI's MASK is a very particular thing: a batch of float tensors in 0..1, one channel. Your array almost certainly isn't. It might be a uint8 binary out of a threshold, a float probability map out of a segmenter, a uint16 depth map, a bool array, or a 0..255 grayscale result from any of the pack's several hundred cv2_* wrappers.

This is the conversion that lands it in the right shape and the right range. If you've done the OpenCV work and now you need an inpaint, a mask composite, or anything else that takes a MASK socket, this is the last node in the chain. (CV Mask → CV is the reverse door, if you're heading the other way.)

How it works

One input: nparray, a single-channel [H,W] array - or a multi-channel one, which gets converted to grayscale first. Any dtype. One output: mask.

The dtype handling is the interesting part, and it's more thoughtful than a cast:

  • uint8 passes straight through and gets divided by 255, so a 0/255 binary mask arrives as 0/1, which is exactly what the mask convention wants.
  • uint16 is divided by 257, landing back on the 0..255 scale before that.
  • bool becomes 0 or 255.
  • Float arrays are rescaled from their own range into 0..255 before the divide. That's what lets a distance transform, a disparity map or a classifier's score map become a visible, usable mask rather than a black square - the values were never in 0..1 to begin with.
  • Non-finite values are clamped to the finite range rather than left to poison it. A stray nan or inf - which distance transforms and optical flow produce routinely - would otherwise make the min/max computation useless and collapse the whole mask to near-black.

A subtlety worth knowing, because it's the difference between "works" and "why is my mask soft": this is a scale, not a threshold. Anything you feed in comes out as a continuous field in proportion to its own values. If you wanted a hard, binary edge, threshold it first (cv2_threshold or one of the pack's threshold nodes) and convert the result.

Where it fits

The productive pattern is: cv2 operation → arithmetic to get the field you want → convert → mask socket. Example: run CV Connected Components or CV Find Contours, build a filled region as a mask array, convert, and feed a per-region inpaint. Another: cv2_distanceTransform → convert → use the soft field as a feathered blend weight, which is a genuinely nice trick and free of any model.

The pack's masking corner has a whole family in both directions - CV Masks to BBoxes, CV BBoxes to Masks, CV Overlay Masks, CV Labels to Masks, CV Keep Largest Component - so this node rarely travels alone.

Install

ComfyUI Manager, search ComfyUI CV. Or:

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv

Restart afterwards. Python ≥ 3.12, a ComfyUI recent enough for the V3 node API, and:

pip install "opencv-contrib-python-headless~=5.0.0.93"

Where people get burned

Min-max scaling makes frames inconsistent. Because float input is scaled from its own per-array range, array A with values 0..0.3 and array B with 0..0.9 both come out spanning the full mask range, at different scales. If you're processing a sequence, normalise explicitly before this node so every frame is on the same scale.

Your multi-channel array went through a BGR assumption. Three or more channels get converted to grayscale as though they were BGR, then reduced. Feeding a heatmap, a LAB array or a three-channel feature map in gives you a mask derived from a channel-weighting you didn't choose. Convert to the single channel you actually mean, upstream, with cv2_extractChannel or a raw cvtColor.

Off-by-255 thinking. A value of 255 in your array is full mask; 255 is not "a large float". If you've been doing arithmetic that produced values up in the thousands, they'll all clamp to full mask and you'll get a solid rectangle. Check with Inspect CV Data when a mask comes out as one flat tone.

It takes an array, not a LATENT. Latents aren't images; if you want a mask out of latent-space data, unwrap it first and think about whether the cell-versus-pixel units still mean what you think they mean.

One array per call. It's a single-frame converter, not a batch one. Looping over a batch is a graph-design question, and the pack's answer to that is usually the list-versus-batch policy idiom rather than running this node a hundred times.

Categoryimage/CV/low-level

Inputs (1)

NameTypeDefaultDescription
nparrayNPARRAYSingle-channel [H,W] ndarray (multi-channel is converted to grayscale first). Any dtype; scaled to a 0-1 mask.

Outputs (1)

NameTypeDescription
maskMASK—