Nodes/ComfyUI CV/CV Scale BBoxes
ComfyUI Node

CV Scale BBoxes

When the box you found is in the wrong coordinate space

By bmad4ever·Created 4 months ago·Updated 14 days ago· 1
CV Scale BBoxes
  • bboxes
  • from_space
  • to_space
  • bboxes
  • scale_x
  • scale_y

Detection nodes hand you boxes. The problem is that they hand them to you in whatever space they were looking at, and that is almost never the space you want to draw or crop in. A template match run on a LATENT reports latent-cell coordinates. A detector run on a 512px downscale reports 512px-space coordinates. CV Scale BBoxes is the one-node fix: give it the boxes, tell it where they came from and where they're going, get boxes back that land on the right pixels.

How it works

It reads exactly two numbers - the height and width of each reference - and computes to_space size / from_space size per axis. That's the whole mechanism. Because it only ever asks the references how big they are, either side can be an IMAGE, a MASK, an NPARRAY or a LATENT. A LATENT's "size" is its latent-cell grid, which is why a stride-8 VAE gives you a factor of 8.0 on both axes without you doing any arithmetic.

It scales the box edges, not the width and height fields - x1 and x2 get scaled independently and the width is re-derived - so two boxes that were touching before are still touching after. It also preserves keys it doesn't understand, so a score or label your detector attached survives the trip.

What you actually set

Three required inputs, and only the middle two need thought:

  • bboxes - core BOUNDING_BOX data. A single dict, a flat list, or per-frame nested lists all work. That's core's own type (the same one PrimitiveBoundingBox emits), so this node sits happily between a core detector and core crop/draw nodes.
  • from_space - the reference the boxes were measured on. The tooltip's own example: the LATENT that was searched. Its tooltip for the output confirms the number you should expect: 8.0 for a stride-8 VAE latent.
  • to_space - the reference you want to land on: the original IMAGE, or its MASK. Again, only height/width are read.

Outputs are bboxes (same nesting as the input, coordinates in to_space), plus scale_x and scale_y as plain floats. Those two floats are more useful than they look - wire them into a CV Scale Points or CV Scale Homography when you have keypoints or a warp matrix from the same detection pass and want everything in one consistent space.

Where it goes in a graph

The canonical chain: downscale, detect, scale back up, crop. Wire the scaled bboxes into core Draw BBoxes to see what you got, or into CV Crop by BBoxes for the actual cut-out. If you're cropping, look at CV Snap BBoxes next to this one - it authors the sizing policy (padding, squareness, VAE-stride multiple) that the crop nodes consume, and the two are designed to be used together.

Install

The pack is ComfyUI CV (bmad4ever/comfyui_cv). In ComfyUI Manager, search the pack title and install, then restart. By hand:

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"

It needs Python ≥ 3.12 and a recent ComfyUI built on the V3 node API. Behaviour is pinned against opencv-contrib 5.0.0.93; other versions may behave differently. No model files needed for this node.

Common issues

  • Looks like a no-op. If both references end up the same size, you get the same boxes back and factors of 1.0. That usually means you wired the same image to both sides - or wired the latent where the image should go.
  • Boxes off by roughly 8x. Almost always the latent/stride factor you told yourself you'd remember. Check scale_x in the node's output rather than eyeballing the preview.
  • "from_space has a degenerate size". The node raises on a zero-sized reference instead of silently returning nonsense, which is the behaviour you want. An empty mask with no dimensions is the usual cause.
  • All the CV nodes vanished after a pip install. You installed a non-contrib OpenCV wheel over the contrib one. All four distributions share one site-packages/cv2, so pip install opencv-python quietly empties the contrib submodules. The repo ships tools/repair_opencv_contrib.py - run it with --check, then --apply.
Categoryimage/CV

Inputs (3)

NameTypeDefaultDescription
bboxesBOUNDING_BOX[object Object]Core BOUNDING_BOX data: per-frame lists of {x, y, width, height} dicts - compatible with Draw BBoxes, Crop By Bounding Boxes, Image Crop, etc. Boxes in from_space coordinates: a single dict, a flat list, or per-frame nested lists. Extra keys (score, label, ...) are preserved.
from_spaceNPARRAY,IMAGE,MASK,LATENTReference the boxes were measured on (e.g. the LATENT that was searched). Only its height/width are used - for a LATENT that is the latent-cell grid size.
to_spaceNPARRAY,IMAGE,MASK,LATENTReference to map the boxes onto (e.g. the original IMAGE, or its MASK - only height/width are used).

Outputs (3)

NameTypeDescription
bboxesBOUNDING_BOXCore BOUNDING_BOX data: per-frame lists of {x, y, width, height} dicts - compatible with Draw BBoxes, Crop By Bounding Boxes, Image Crop, etc. Same nesting as the input, coordinates in to_space.
scale_xFLOATApplied horizontal factor: to_space width / from_space width (8.0 for a stride-8 VAE latent).
scale_yFLOATApplied vertical factor: to_space height / from_space height.