CV Scale BBoxes
When the box you found is in the wrong coordinate space
- bboxes
- from_space
- to_space
- bboxes
- scale_x
- scale_y
Detection nodes hand you boxes. The problem is that they hand them to you in whatever space they were looking at, and that is almost never the space you want to draw or crop in. A template match run on a LATENT reports latent-cell coordinates. A detector run on a 512px downscale reports 512px-space coordinates. CV Scale BBoxes is the one-node fix: give it the boxes, tell it where they came from and where they're going, get boxes back that land on the right pixels.
How it works
It reads exactly two numbers - the height and width of each reference - and computes to_space size / from_space size per axis. That's the whole mechanism. Because it only ever asks the references how big they are, either side can be an IMAGE, a MASK, an NPARRAY or a LATENT. A LATENT's "size" is its latent-cell grid, which is why a stride-8 VAE gives you a factor of 8.0 on both axes without you doing any arithmetic.
It scales the box edges, not the width and height fields - x1 and x2 get scaled independently and the width is re-derived - so two boxes that were touching before are still touching after. It also preserves keys it doesn't understand, so a score or label your detector attached survives the trip.
What you actually set
Three required inputs, and only the middle two need thought:
- bboxes - core
BOUNDING_BOXdata. A single dict, a flat list, or per-frame nested lists all work. That's core's own type (the same onePrimitiveBoundingBoxemits), so this node sits happily between a core detector and core crop/draw nodes. - from_space - the reference the boxes were measured on. The tooltip's own example: the LATENT that was searched. Its tooltip for the output confirms the number you should expect: 8.0 for a stride-8 VAE latent.
- to_space - the reference you want to land on: the original IMAGE, or its MASK. Again, only height/width are read.
Outputs are bboxes (same nesting as the input, coordinates in to_space), plus scale_x and scale_y as plain floats. Those two floats are more useful than they look - wire them into a CV Scale Points or CV Scale Homography when you have keypoints or a warp matrix from the same detection pass and want everything in one consistent space.
Where it goes in a graph
The canonical chain: downscale, detect, scale back up, crop. Wire the scaled bboxes into core Draw BBoxes to see what you got, or into CV Crop by BBoxes for the actual cut-out. If you're cropping, look at CV Snap BBoxes next to this one - it authors the sizing policy (padding, squareness, VAE-stride multiple) that the crop nodes consume, and the two are designed to be used together.
Install
The pack is ComfyUI CV (bmad4ever/comfyui_cv). In ComfyUI Manager, search the pack title and install, then restart. By hand:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"
It needs Python ≥ 3.12 and a recent ComfyUI built on the V3 node API. Behaviour is pinned against opencv-contrib 5.0.0.93; other versions may behave differently. No model files needed for this node.
Common issues
- Looks like a no-op. If both references end up the same size, you get the same boxes back and factors of 1.0. That usually means you wired the same image to both sides - or wired the latent where the image should go.
- Boxes off by roughly 8x. Almost always the latent/stride factor you told yourself you'd remember. Check
scale_xin the node's output rather than eyeballing the preview. - "from_space has a degenerate size". The node raises on a zero-sized reference instead of silently returning nonsense, which is the behaviour you want. An empty mask with no dimensions is the usual cause.
- All the CV nodes vanished after a
pip install. You installed a non-contrib OpenCV wheel over the contrib one. All four distributions share onesite-packages/cv2, sopip install opencv-pythonquietly empties the contrib submodules. The repo shipstools/repair_opencv_contrib.py- run it with--check, then--apply.
Inputs (3)
| Name | Type | Default | Description |
|---|---|---|---|
| bboxes | BOUNDING_BOX | [object Object] | Core BOUNDING_BOX data: per-frame lists of {x, y, width, height} dicts - compatible with Draw BBoxes, Crop By Bounding Boxes, Image Crop, etc. Boxes in from_space coordinates: a single dict, a flat list, or per-frame nested lists. Extra keys (score, label, ...) are preserved. |
| from_space | NPARRAY,IMAGE,MASK,LATENT | Reference the boxes were measured on (e.g. the LATENT that was searched). Only its height/width are used - for a LATENT that is the latent-cell grid size. | |
| to_space | NPARRAY,IMAGE,MASK,LATENT | Reference to map the boxes onto (e.g. the original IMAGE, or its MASK - only height/width are used). |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| bboxes | BOUNDING_BOX | Core BOUNDING_BOX data: per-frame lists of {x, y, width, height} dicts - compatible with Draw BBoxes, Crop By Bounding Boxes, Image Crop, etc. Same nesting as the input, coordinates in to_space. |
| scale_x | FLOAT | Applied horizontal factor: to_space width / from_space width (8.0 for a stride-8 VAE latent). |
| scale_y | FLOAT | Applied vertical factor: to_space height / from_space height. |