CV Array Shape
Stop guessing what a cv2 node just handed you
- nparray
- height
- width
- channels
- dsize
- size
What it's for
Half the friction of doing computer vision inside ComfyUI is shape bookkeeping. You run a cv2 wrapper, get an array back, and now need to know its height, width and channel count - to size a dsize, to build a kernel, to write a mask, or just to figure out why the next node complained. There is no print statement in a node graph.
CV Array Shape is the print statement. It reads the spatial dimensions and channel count of an IMAGE, MASK, LATENT or raw NPARRAY and hands them back as integers - no data read, no copy, no modification. It's the shape half of the debugging pair, with Inspect CV Data doing the values.
It also makes the pack's low-level cv2_* wrappers usable without a spreadsheet: those take composite parameters (dsize, ksize, winSize, imageSize) that people hard-code until the input resolution changes.
The outputs
Four, and the naming is the whole point:
- height - first spatial dimension.
- width - second spatial dimension.
- channels - 1 for a 2-D array or a
MASK;Cfor an[H,W,C]array, anIMAGEor aLATENT. - dsize - the
(width, height)pair as a literal string, for the composite params that are still free text. - size - the same pair as a single
CV_TUPLEvalue, which you wire straight into a wrapper'sdsize/ksize/winSize/imageSizeinput.
Use size and not dsize whenever the input is type-aware. Both components travel on one wire, so the size can't arrive half-connected - with the string version you're one typo away from feeding "(1024, 768)" into something that wanted it the other way round. That other way round is the classic: OpenCV's dsize is (width, height), most people's intuition is (height, width), and a swapped resize is not a crash, it's a squashed image you might not notice for a while.
How it behaves
The input accepts NPARRAY, IMAGE, MASK or LATENT. An IMAGE is read as frame 0 of the batch - this is a shape reader for one frame, not a batch summariser. Only the shape is inspected regardless of dtype, so it works on anything from a uint8 mask to a float64 matrix.
For a LATENT the numbers count latent cells, not pixels - pixels divided by the VAE stride, which is 8 for the usual autoencoder. A 1024×1024 SDXL latent reports 128×128. That isn't a quirk, it's the units every cv2 node downstream is actually working in when you hand it a latent.
Where it fits
Shape-driven graphs usually look like this: decode or load something, read size off it once, and fan that one CV_TUPLE out to every cv2_resize, cv2_getStructuringElement and filter that needs to agree. Change the source resolution and the graph follows. Wire width and height separately when you need arithmetic rather than a pair - cropping math, grid counts, area checks.
The sibling node is CV Array Size in the features category, which gives you width, height, dsize and size but not channels. Reach for that one when you're checking dimensions; reach for this one the moment channel count matters, which in this pack is often, because so many raw wrappers care whether they got grey or colour.
Install
ComfyUI Manager, search ComfyUI CV. Or:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
Then restart, and keep the contrib wheel in place:
pip install "opencv-contrib-python-headless~=5.0.0.93"
Python ≥ 3.12, plus a ComfyUI recent enough for the V3 node API.
Where people get burned
Latent cells read as pixels. You check a latent, see 128, and build a 128-pixel kernel for it. At latent scale that kernel is 8× bigger than you meant. Multiply by the stride when you're thinking in image pixels.
Frame 0 only. Wire a 40-frame batch in and you get the shape of the first frame. If your frames genuinely differ in size, that's a problem this node can't describe and probably one you want to fix upstream.
Two wires, one input. dsize and size are alternatives, not a pair - connecting both to the same socket is not how any of these nodes work. Pick size.
dsize is width-comma-height. If you're using the string output with an old free-text parameter, that's the order, parentheses included, exactly as printed.
And the pack-level caveat, since it bites this node category: the ~470 cv2_* wrappers are generated from whatever OpenCV your install exposes, so a stripped-down wheel can make a node you saw in a screenshot simply not exist - and the raw wrappers are explicitly uncurated. Expect to handle units and dtypes yourself. This node is how you find out what you actually got.
Inputs (1)
| Name | Type | Default | Description |
|---|---|---|---|
| nparray | NPARRAY,IMAGE,MASK,LATENT | Array whose shape to read. Only the shape is inspected regardless of dtype. LATENT sizes count latent cells. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. |
Outputs (5)
| Name | Type | Description |
|---|---|---|
| height | INT | Height (first spatial dimension). For LATENT inputs this counts latent cells (pixels / stride). |
| width | INT | Width (second spatial dimension). For LATENT inputs this counts latent cells. |
| channels | INT | Number of channels: 1 for 2-D arrays / MASK, C for [H,W,C] arrays / IMAGE / LATENT. |
| dsize | STRING | '(width, height)' literal, for the composite cv2 params that are still free-text. Prefer the 'size' output for a Size-typed input. |
| size | CV_TUPLE | (width, height) as ONE composite value - wire it straight into a cv2 node's dsize / ksize / winSize / imageSize input. Both components travel together, so the size can never arrive half-connected. |