CV Array Size
One wire for (width, height), no more swapped resizes
- nparray
- width
- height
- dsize
- size
What it's for
OpenCV's dsize is (width, height). Nearly everyone's instinct is (height, width). That mismatch is one of the most productive bug generators in all of computer vision, and inside a node graph it's worse than in code, because there's nowhere to print the value and check.
This node reads the width and height of an image array and hands them back - and, crucially, hands the pair back as one composite value (size, a CV_TUPLE) so you can wire it into a wrapper's dsize / ksize / winSize / imageSize input in a single link. Both components travel together. The size can't arrive half-connected the way two loose integer wires can, and it can't arrive in the wrong order because the pair was built by the node that read the shape.
There's also a dsize output: the same (width, height) pair as a literal string, for the composite parameters in the low-level wrappers that are still free text. If a parameter accepts a typed size, prefer size.
How it works
The input takes NPARRAY, IMAGE, MASK or LATENT, and only the shape is used - nothing is read, copied or modified. An IMAGE is read as frame 0 of the batch.
For a LATENT the numbers are in latent cells, not pixels: pixels divided by the VAE stride, 8 for the usual autoencoder. So a 1024×1024 image's latent reports 128×128. That's the right answer for the cv2 nodes, which work on the latent as it is, and the wrong answer for anything that wants to know how big the decoded picture will be. Know which one you're holding.
Where it fits
The reason to bother is fan-out. Wire one CV Array Size into every resize, every filter kernel, every morphology structuring element in the branch, and the whole thing follows your input resolution instead of pinning numbers you typed while testing on one image.
That's the same discipline as the seed primitive: one authoritative source, many consumers, no silent mismatches. In this pack it matters more than usual, because so many parameters are composite - dsize, ksize, winSize, imageSize - and a composite that arrives as two independent integers can be half-connected (one wire forgotten) or swapped, both of which produce plausible-looking output.
The sibling to reach for when you also care about channel count is CV Array Shape, which adds a channels output (and a height/width naming that's identical otherwise). CV Array Size lives in the pack's features category; CV Array Shape sits with the low-level IO nodes.
Install
ComfyUI Manager, search ComfyUI CV, or:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
Restart. Requirements: Python ≥ 3.12, a ComfyUI recent enough to have the V3 node API, and:
pip install "opencv-contrib-python-headless~=5.0.0.93"
Where people get burned
Latent cells read as pixels. The most common one. You check a latent, get 128, and size a kernel or a grid for it as if that were the picture size - which is 8× off in each direction. It's not wrong output, it's the correct output in the wrong units.
Feeding both outputs to the same socket. size (the composite) and dsize (the string) are alternatives, not a pair. Pick one.
size is (width, height), and so is cv2's. The whole value of this node is that it gets that right where a human wouldn't - but it only helps if you use the composite instead of reading width and height separately and rebuilding the pair by hand. The moment you retype it, you've reintroduced the bug the node exists to remove.
A batch gives you frame 0. This is a per-frame reader. If your frames differ in size, this node won't tell you, and you probably want to fix that upstream rather than downstream.
Resolution changes mid-graph. If a resize happens after your CV Array Size node, everything downstream of the size read is describing the old shape. Read the size after the step that changes it - or better, after every step that changes it. The node is cheap; confusion is not.
Inputs (1)
| Name | Type | Default | Description |
|---|---|---|---|
| nparray | NPARRAY,IMAGE,MASK,LATENT | Image whose width/height to read (only the shape is used). LATENT sizes count latent cells, not pixels. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. |
Outputs (4)
| Name | Type | Description |
|---|---|---|
| width | INT | — |
| height | INT | — |
| dsize | STRING | '(width, height)' literal - for the composite cv2 params that are still free-text. |
| size | CV_TUPLE | (width, height) as ONE composite value - wire it straight into a cv2 node's dsize / ksize / imageSize input, in one link instead of two. |