CV DNN Blob From Image
The preprocessing step every ONNX model needs and nobody documents
- image
- blob
CV DNN Blob From Image builds the NCHW input tensor an ONNX network expects: an optional resize, per-channel mean subtraction, a scale factor and a BGR↔RGB swap, packed into one float32 blob.
This is the step that nobody writes tutorials about and that breaks most "I loaded an ONNX model and it produced garbage" attempts. Networks are picky. They want a specific square size, a specific channel order, and pixels in the range they were trained on. Get the mean wrong and you get plausible-looking noise; get the channel order wrong and you get subtle colour inversion that you'll blame on the model.
Pair it with CV DNN Forward (or CV DNN Forward All) and CV DNN Images From Blob and you have the generic three-step pipeline for any ONNX image model: preprocess, infer, postprocess.
How it works
cv2.dnn.blobFromImage under the hood, over a whole batch at once. It resizes to the width/height you ask for (0 means keep the image's own dimension, which is what fully-convolutional models like flow nets need), subtracts the means, multiplies by scale, optionally swaps R and B, and lays the result out as (N, C, H, W) - channels-first, unit float.
The softmax/normalisation options that the cv2 function offers aren't exposed, and that's deliberate: modern ONNX exports fold normalisation into the graph, so exposing it mostly creates ways to double-apply it.
The inputs that matter
Required:
image- an IMAGE batch (every frame is encoded) or a single NPARRAY, which is treated as one BGR frame whatever its dtype.scale- default 1, meaning "leave pixels at 0–255". Set0.003921(1/255) to map them into 0–1, which is what most ONNX models expect.width/height- the network's input size.swap_rb- leave it on. OpenCV frames are BGR; nearly every model out there was trained on RGB.
Optional: mean_r, mean_g, mean_b (all 0 = no subtraction; set the model's training means if it has them) and crop (off by default - turns the resize into "fit the shorter side, then centre-crop", which is what aspect-preserving detection models want).
Output: blob, an (N, C, H, W) float32 array. An empty input batch gives you the strange-looking (0, 3, 1, 1) rather than an error, so an unexpected shape here usually means the image never arrived.
Install
Part of comfyui_cv (bmad4ever/comfyui_cv). ComfyUI Manager → search "ComfyUI CV", or:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"
Restart ComfyUI. Python ≥ 3.12 and a recent ComfyUI on the V3 node API.
Then the part the README is emphatic about: models are not bundled. Drop your .onnx files in ComfyUI/models/onnx - that's where the model dropdown reads from - and look at the pack's model_sources.txt for the URLs, target folders and licences of the models its workflows use. None of them ship with the repo, and a couple carry licences stricter than the pack's own GPL.
Where people get burned
- The output image is inside out, colour-wise.
swap_rboff, or an unusually BGR-trained model. Flip it and look again. - The model's output is noise. Nine times out of ten it's
scale(you left it at 1 when the model wants 0–1) or the means. Check the model card; there's no way to infer this from the graph reliably. - Output is stretched or squashed. Aspect-changing resize. Turn on
crop, or setwidth/heightto 0 and let the model take its native resolution if it's fully convolutional. - Everything is slow. This is
cv2.dnnon the CPU, and the pack's README is completely straightforward about the tradeoff: ComfyUI core often does this job better. The clearest example is frame interpolation - core ships a native RIFE/FILM node that runs on the GPU in fp16 with offloading and a batch multiplier, while the pack's example does an ONNX export through these nodes on the CPU, one frame pair per execution. Use this chain to understand the pieces or to reach a model core doesn't support. Don't use it because it's there. - Half of your batch vanished. The batch dimension is
N; a model that froze batch 1 into its constants (many two-input exports do) will fail or return one result. Feed one at a time.
Inputs (9)
| Name | Type | Default | Description |
|---|---|---|---|
| image | NPARRAY,IMAGE | IMAGE batch to encode. Every frame is resized and encoded; an NPARRAY (gray/BGR/BGRA, any dtype) is treated as a single BGR frame. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| scale | FLOAT | 1.0000–1000000 | Multiplies every pixel (after mean subtraction). Use 1/255 = 0.003921 to map 0-255 pixels into the 0-1 range most ONNX models expect; 1.0 leaves them at 0-255. |
| width | INT | 00–8192 | Resize width the network wants, in px. 0 = keep the image's own width (for fully-convolutional models). |
| height | INT | 00–8192 | Resize height the network wants, in px. 0 = keep the image's own height. |
| swap_rb | BOOLEAN | true | Swap the R and B channels. OpenCV frames are BGR but most ONNX models are trained on RGB, so leave this on. |
| mean_ropt | FLOAT | 0.0-255–255 | Value subtracted from the RED channel before scaling (0 = none). Set the model's training mean if needed. |
| mean_gopt | FLOAT | 0.0-255–255 | Value subtracted from the GREEN channel before scaling (0 = none). |
| mean_bopt | FLOAT | 0.0-255–255 | Value subtracted from the BLUE channel before scaling (0 = none). |
| cropopt | BOOLEAN | false | On: resize so the shorter side fits then centre-crop to (width, height). Off: resize the whole image (may change aspect ratio). |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| blob | NPARRAY | (N, C, H, W) float32 NCHW blob for the whole batch. Feed 'CV DNN Forward'. Empty (0, 3, 1, 1) for an empty IMAGE batch. |