cv2.dnn.blobFromImage
What a neural net actually wants as input
- image
- size
- nparray
The step everyone forgets when they hand-roll ONNX inference: a network doesn't want a picture, it wants a blob - a 4-D NCHW float tensor, channel order fixed, size fixed, mean subtracted, pixels scaled. cv2.dnn.blobFromImage is that step, and it crams resize, crop, mean subtraction, scaling and the BGR→RGB swap into one call so you don't have to remember the order.
The mechanics, because the knobs only make sense together
The output is (1, C, H, W) - one batch entry, channels first. That shape is the thing; everything else is how the pixels in it got computed.
imageis the input, and here it's aNPARRAY,IMAGEorMASK. An IMAGE comes through as uint8 BGR, a MASK as single-channel - andblobFromImagegenuinely accepts 1-, 3- and 4-channel inputs, so a grayscale blob (1, 1, H, W) is a legitimate thing to ask for.sizeis aCV_TUPLE(width, height), default(0, 0), which means "keep the image's own size". That default is why your first blob can look right while a model trained at 224×224 immediately rejects it.scalefactor,swapRBandcropare blank-by-default string widgets that take Python literals -"1.0","true"- and blank means "don't pass the argument, use OpenCV's default". SoswapRBblank isfalse, which with this pack's BGR-ordered input means your model gets BGR unless you type"true". Half the ONNX models in the wild were trained on RGB; this is where they quietly lose accuracy.cropis what makessizebehave. Withcrop = falseand a non-matching aspect ratio, the image is squashed to fit. Withcrop = true, it's resized so the short side fits and then centre-cropped - which is what most fixed-size classifiers actually expect.meanis acv2Scalar written as a literal,"(104, 117, 123)"style, subtracted after the scale. It's the mean your model's author used. Guessing it wrong is a classic accuracy leak that looks like a bad model.ddepthis the output depth dropdown. Leave it on "same as input"; OpenCV's own docs say chooseCV_32ForCV_8U, and the pack's default is the-1sentinel.
The single output is nparray: the blob. It's not an image and there's no format echo here - it will not preview.
Where it sits in a graph
It's the front half of the generic cv2.dnn pipeline: blobFromImage → CV DNN Forward (which loads an ONNX and pushes the blob through), then whatever decoder matches the model. The pack's curated CV DNN Blob From Image node does the batch version of the same job with typed options, which is nicer if you have more than one frame; this raw wrapper takes frame 0.
One caveat the pack's README states better than I can: this whole DNN corner is deliberately CPU-side ONNX, and ComfyUI often has a first-class node for the same task that runs faster in PyTorch. Use this when you're reaching a model ComfyUI doesn't support, or to understand how the pieces fit.
Installing comfyui_cv
Part of bmad4ever/comfyui_cv - ~470 auto-generated raw cv2.* wrappers plus curated nodes, a GPL-3.0 fork of Gerold Meisinger's opencv-comfyui. Manager: search ComfyUI CV. Or:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
Restart. Python ≥ 3.12 and a recent ComfyUI on the V3 node API; dependency is opencv-contrib-python-headless~=5.0.0.93 - the contrib wheel specifically, since the four OpenCV distributions share one site-packages/cv2 and a non-contrib install last-wins the contrib modules away. Models are not bundled; the pack's model_sources.txt records the URLs and licences.
Common issues
The model errors on shape. Your size is (0, 0) or doesn't match the network's expected input. Check the model's input layer and set both components.
Accuracy is mysteriously poor. Nine times out of ten it's swapRB (the pack feeds BGR, the model may want RGB) or a wrong mean. Fix those two before you blame the ONNX export.
Squashed content. crop is false with a mismatched aspect ratio. Turn it on.
"Can't convert to np.ndarray" style errors from the next node. You wired the blob into an image input. The blob is data; it goes to something that forwards it.
Blank widget not working. Type a real literal. Blank means "omit the argument", which is not the same as typing 0.
Inputs (7)
| Name | Type | Default | Description |
|---|---|---|---|
| image | NPARRAY,IMAGE,MASK | - - - Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| scalefactoropt | STRING | - - - Optional - leave blank to use the OpenCV default. Accepts a Python literal, e.g. 3, 1.5, true, or (3, 3). | |
| sizeopt | CV_TUPLE | 0,0 | One value with 2 components (w, h) - it travels as a whole, so it cannot arrive half-connected. Wire it from 'CV Tuple' or type the components in place. |
| meanopt | STRING | - - - cv2 Scalar as a literal, e.g. "(0, 255, 0)" (BGR) or "(0, 255, 0, 64)" (BGRA). A bare number broadcasts to every component, so "255" means (255, 255, 255, 255). Components past the target's channel count are ignored by OpenCV. Leave blank for the OpenCV default. | |
| swapRBopt | STRING | - - - Optional - leave blank to use the OpenCV default. Accepts a Python literal, e.g. 3, 1.5, true, or (3, 3). | |
| cropopt | STRING | - - - Optional - leave blank to use the OpenCV default. Accepts a Python literal, e.g. 3, 1.5, true, or (3, 3). | |
| ddepthopt | COMBO | same as input | - - - |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| nparray | NPARRAY | — |