Nodes/ComfyUI CV/CV DNN Blob From Image
ComfyUI Node

CV DNN Blob From Image

The preprocessing step every ONNX model needs and nobody documents

By bmad4ever·Created 4 months ago·Updated 15 days ago· 1
CV DNN Blob From Image
  • image
  • blob
◄scale1.000►
◄width0►
◄height0►
◄swap_rbtrue►
◄mean_r0.0►
◄mean_g0.0►
◄mean_b0.0►
◄cropfalse►

CV DNN Blob From Image builds the NCHW input tensor an ONNX network expects: an optional resize, per-channel mean subtraction, a scale factor and a BGR↔RGB swap, packed into one float32 blob.

This is the step that nobody writes tutorials about and that breaks most "I loaded an ONNX model and it produced garbage" attempts. Networks are picky. They want a specific square size, a specific channel order, and pixels in the range they were trained on. Get the mean wrong and you get plausible-looking noise; get the channel order wrong and you get subtle colour inversion that you'll blame on the model.

Pair it with CV DNN Forward (or CV DNN Forward All) and CV DNN Images From Blob and you have the generic three-step pipeline for any ONNX image model: preprocess, infer, postprocess.

How it works

cv2.dnn.blobFromImage under the hood, over a whole batch at once. It resizes to the width/height you ask for (0 means keep the image's own dimension, which is what fully-convolutional models like flow nets need), subtracts the means, multiplies by scale, optionally swaps R and B, and lays the result out as (N, C, H, W) - channels-first, unit float.

The softmax/normalisation options that the cv2 function offers aren't exposed, and that's deliberate: modern ONNX exports fold normalisation into the graph, so exposing it mostly creates ways to double-apply it.

The inputs that matter

Required:

  • image - an IMAGE batch (every frame is encoded) or a single NPARRAY, which is treated as one BGR frame whatever its dtype.
  • scale - default 1, meaning "leave pixels at 0–255". Set 0.003921 (1/255) to map them into 0–1, which is what most ONNX models expect.
  • width / height - the network's input size.
  • swap_rb - leave it on. OpenCV frames are BGR; nearly every model out there was trained on RGB.

Optional: mean_r, mean_g, mean_b (all 0 = no subtraction; set the model's training means if it has them) and crop (off by default - turns the resize into "fit the shorter side, then centre-crop", which is what aspect-preserving detection models want).

Output: blob, an (N, C, H, W) float32 array. An empty input batch gives you the strange-looking (0, 3, 1, 1) rather than an error, so an unexpected shape here usually means the image never arrived.

Install

Part of comfyui_cv (bmad4ever/comfyui_cv). ComfyUI Manager → search "ComfyUI CV", or:

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"

Restart ComfyUI. Python ≥ 3.12 and a recent ComfyUI on the V3 node API.

Then the part the README is emphatic about: models are not bundled. Drop your .onnx files in ComfyUI/models/onnx - that's where the model dropdown reads from - and look at the pack's model_sources.txt for the URLs, target folders and licences of the models its workflows use. None of them ship with the repo, and a couple carry licences stricter than the pack's own GPL.

Where people get burned

  • The output image is inside out, colour-wise. swap_rb off, or an unusually BGR-trained model. Flip it and look again.
  • The model's output is noise. Nine times out of ten it's scale (you left it at 1 when the model wants 0–1) or the means. Check the model card; there's no way to infer this from the graph reliably.
  • Output is stretched or squashed. Aspect-changing resize. Turn on crop, or set width/height to 0 and let the model take its native resolution if it's fully convolutional.
  • Everything is slow. This is cv2.dnn on the CPU, and the pack's README is completely straightforward about the tradeoff: ComfyUI core often does this job better. The clearest example is frame interpolation - core ships a native RIFE/FILM node that runs on the GPU in fp16 with offloading and a batch multiplier, while the pack's example does an ONNX export through these nodes on the CPU, one frame pair per execution. Use this chain to understand the pieces or to reach a model core doesn't support. Don't use it because it's there.
  • Half of your batch vanished. The batch dimension is N; a model that froze batch 1 into its constants (many two-input exports do) will fail or return one result. Feed one at a time.
Categoryimage/CV/dnn

Inputs (9)

NameTypeDefaultDescription
imageNPARRAY,IMAGEIMAGE batch to encode. Every frame is resized and encoded; an NPARRAY (gray/BGR/BGRA, any dtype) is treated as a single BGR frame. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
scaleFLOAT1.0000–1000000Multiplies every pixel (after mean subtraction). Use 1/255 = 0.003921 to map 0-255 pixels into the 0-1 range most ONNX models expect; 1.0 leaves them at 0-255.
widthINT00–8192Resize width the network wants, in px. 0 = keep the image's own width (for fully-convolutional models).
heightINT00–8192Resize height the network wants, in px. 0 = keep the image's own height.
swap_rbBOOLEANtrueSwap the R and B channels. OpenCV frames are BGR but most ONNX models are trained on RGB, so leave this on.
mean_roptFLOAT0.0-255–255Value subtracted from the RED channel before scaling (0 = none). Set the model's training mean if needed.
mean_goptFLOAT0.0-255–255Value subtracted from the GREEN channel before scaling (0 = none).
mean_boptFLOAT0.0-255–255Value subtracted from the BLUE channel before scaling (0 = none).
cropoptBOOLEANfalseOn: resize so the shorter side fits then centre-crop to (width, height). Off: resize the whole image (may change aspect ratio).

Outputs (1)

NameTypeDescription
blobNPARRAY(N, C, H, W) float32 NCHW blob for the whole batch. Feed 'CV DNN Forward'. Empty (0, 3, 1, 1) for an empty IMAGE batch.