Nodes/ComfyUI CV/CV DNN Letterbox
ComfyUI Node

CV DNN Letterbox

The boring resize that your YOLO graph silently depends on

By bmad4ever·Created 3 months ago·Updated 14 days ago· 1
CV DNN Letterbox
  • image
  • image
  • ratio
  • pad_left
  • pad_top
◄size640►
◄pad_value114►

Why you'd reach for this

Every YOLO-family ONNX model has an opinion about its input shape: one fixed square, usually 640×640, and nothing else. Your photo is a 16:9 landscape. So you resize it without distorting it, then fill the leftover strip with flat grey. That is a letterbox.

It's the least interesting node in the pipeline and one of the two that decide whether your detection boxes are right or garbage. The model sees the padded square, so it reports boxes in padded coordinates. Undo the scale and the padding wrong and every box slides toward the top-left corner - the classic "why is my YOLO detecting thing off by 40 pixels" bug. This node hands you back the numbers you need to undo it.

How it works

Aspect-preserving: it computes ratio = min(size / h, size / w), resizes to (round(w*ratio), round(h*ratio)) with linear interpolation, and pads the remainder with cv2.copyMakeBorder in a constant grey. It follows the Ultralytics LetterBox convention down to the rounding - the paddings are computed as round(d - 0.1) and round(d + 0.1) for the two sides of each axis, which is Ultralytics' trick for splitting an odd number of pixels without drifting.

A ComfyUI IMAGE batch is letterboxed frame by frame, but all frames in a batch share one size, so there is exactly one ratio/pad triplet for the whole batch, taken from frame 0. Feed it a single frame or an NPARRAY and it just does the one.

Inputs and outputs that matter

  • size - the model's imgsz, square, default 640, stepped in 32s. Match it to the export you actually downloaded; 640 vs 416 is not a rounding detail, it's a wrong answer.
  • pad_value - the grey, default 114, which is the Ultralytics value. Leave it alone unless the model card says otherwise; a model trained on black padding will put weak detections in the pad strip if you feed it 114.
  • image - IMAGE or NPARRAY in, and note the output is a plain IMAGE, because the next stop is CV DNN Blob From Image, which wants a tensor-ish image, not an ndarray.
  • ratio, pad_left, pad_top - the three numbers the decoder needs. padded_px = orig_px * ratio, so to go back to your original pixels it's orig = (padded - pad_left) / ratio. Wire these into CV YOLO Detect Decode or CV YOLO Seg Masks and they do that arithmetic for you. Only left/top come out, which is all the decode needs because the mapping is uniform.

The usual wiring is a straight line: image → CV DNN Letterbox → CV DNN Blob From Image → CV DNN Forward, then the same letterbox's ratio/pads into the decoder that reads the model's output. Get that line wrong and nothing errors; you just get bad boxes, which is worse.

Install

Same as every node in this pack - one Python dependency, and it has to be the contrib build:

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"

Needs Python ≥ 3.12 and a ComfyUI new enough for the V3 node API. ComfyUI Manager also lists the pack as ComfyUI CV (publisher bmad4ever). Models are not bundled: the .onnx files go in ComfyUI/models/onnx, and the repo's model_sources.txt has the download URLs and a license note per model - worth reading before you redistribute anything, because at least one of the YOLO exports is AGPL-3.0.

Common issues

  • Boxes drift toward the corner. You wired the padded image to a decoder but not the ratio/pads, or you fed the decoder the original ratio. The three scalars and the image have to come from the same letterbox node.
  • Contrib nodes vanish after installing something else. All four OpenCV wheel variants share one site-packages/cv2. Install plain opencv-python over the contrib build and the contrib submodules are emptied - the pack ships tools/repair_opencv_contrib.py --check / --apply for exactly this.
  • Wrong size. A 640 export fed a 416 letterbox will happily run and produce mush. The node can't check this for you.
  • Honest caveat, from the author's own README: this pack deliberately routes everything through cv2.dnn, and for jobs ComfyUI already does natively - frame interpolation is the example the README gives - the core node is faster, GPU-side, fp16, and batched, while the cv2 route is CPU and one item per run. If core has your model, use core. Letterboxing a detection input is the case where there is no core equivalent, so this is genuinely the node to use.
Categoryimage/CV/dnn

Inputs (3)

NameTypeDefaultDescription
imageNPARRAY,IMAGEInput image. An IMAGE batch is letterboxed frame by frame (all frames share one size, hence one ratio/pad); an NPARRAY is treated as a single frame. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
sizeINT64032–4096Square network input size in pixels (the model's imgsz, e.g. 640 for YOLO seg models).
pad_valueoptINT1140–255Grey value used for the letterbox padding. 114 matches the Ultralytics default.

Outputs (4)

NameTypeDescription
imageIMAGEThe letterboxed (size, size) image. Feed 'CV DNN Blob From Image'.
ratioFLOATResize ratio applied (padded_px = orig_px * ratio). Feed the decoder's 'ratio' input.
pad_leftINTLeft padding in letterboxed pixels. Feed the decoder's 'pad_left' input.
pad_topINTTop padding in letterboxed pixels. Feed the decoder's 'pad_top' input.