Nodes/ComfyUI CV/cv2.getRectSubPix
ComfyUI Node

cv2.getRectSubPix

Crop at x=100.5 and get away with it

By bmad4ever·Created 4 months ago·Updated 15 days ago· 1
cv2.getRectSubPix
  • image
  • patchSize
  • center
  • result
◄patchTypesame as input►

Every crop node in ComfyUI takes integers, because pixels are integers. This one takes a floating point centre and interpolates the patch out of the source image, so "the patch centred on the corner at (412.37, 208.91)" is a thing you can actually ask for. That half-pixel sounds like pedantry until you're doing sub-pixel tracking or frame stabilisation, where the answer is supposed to be fractional and rounding it away is the same as throwing away the precision you just paid for.

What it takes and what comes back

  • image - the source. This socket is type-preserving: link an IMAGE or a MASK straight in, or an NPARRAY. CV Array → Image and Image → CV Array move between the two worlds if you need to.
  • patchSize - the extracted patch size (w, h), a composite CV_TUPLE. Integers, because the output has to be a real grid of pixels.
  • center - the (x, y) centre, and this one is genuinely floating point: (412.5, 208.25) is a legal, meaningful value.
  • patchType (optional) - the depth of the extracted pixels, defaulting to same as input. A dropdown of the pack's shared depth enum; leave it alone unless you have a reason to change dtype mid-pipeline.

One output, result, and it echoes whatever you fed the image socket: an IMAGE in comes back as an IMAGE, a MASK as a MASK, an NPARRAY stays an NPARRAY. That's the pack's match-type convention, and it means this node slots into a normal graph without a conversion dance.

The operation itself is a bilinear sample of the patch around that centre - mechanically it's a tiny warpAffine with the interpolation done for you. The center must fall inside the image; for any part of the patch that hangs off the edge, OpenCV replicates the border pixels instead of erroring, which is convenient for edge-tracked objects and worth knowing about before you read a stretched strip as real data.

The batch behaviour, because the tooltips disagree

The pack's generic image-input tooltip says "frame 0 of a batch", and for a lot of wrappers that's exactly right. This one is on the source's per-frame list: link a multi-frame IMAGE and every frame gets its own patch, and the results come back stacked as a batch. Same centre and same size across the sequence - the stabilisation move, essentially: one tracked point, a patch per frame, no resizing the whole image. In shipped workflows you'll more often see it used the simple way, on a single frame, where its advantage over a plain crop is only the fractional centre.

Two limits worth flagging. It's excluded from the pack's latent-array lane deliberately: this is a 1- or 3-channel operation, so don't try to pull patches out of a 4- or 16-channel latent with it - you'll get an error rather than something clever. And it doesn't find anything: the centre is your problem. In practice you feed it the float output of something like CV Phase Correlate (Translation) (a sub-pixel global shift) or a tracked point from CV Track Features (KLT), and this node is the sampling step that consumes it.## Install

pip install "opencv-contrib-python-headless~=5.0.0.93"

ComfyUI Manager → search ComfyUI CV, or:

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv

Restart ComfyUI afterwards. Python ≥ 3.12 and a recent ComfyUI on the V3 node API are required - the pack builds every node at import time, and on an older ComfyUI you get an empty pack rather than a clean error. No models to download; the whole low-level lane is arithmetic over pixels you already have.

Ergonomic annoyances, stated plainly

  • patchSize and center are composite values, so they can't arrive half-connected: author them with CV Tuple or type the (x, y) literal into the widget in place. If your centre comes from a node that emits separate numbers, CV Tuple is the joiner - and its component sockets take INT or float links.
  • The widget defaults are (0, 0) for both, which is not a usable configuration - a 0×0 patch is not "whatever fits". Type real numbers or link a real composite.
  • No preview for free. result is an image-shaped output, but if you fed it an NPARRAY the pack treats it as data, not a picture: Preview CV Array renders raw arrays directly (with normalize/heatmap modes for float data), and Inspect CV Data gives you the shape and dtype when something looks off by a channel.
  • The whole pack is a moving target by design. Its README says updates are unplanned and that the auto-generated wrappers are uncurated - you handle conversions and edge cases. For a four-argument wrapper that's a small risk, but it's the standing posture of this pack, and it's why the curated nodes exist alongside this one.

If you mainly want "the region around a detected box", reach for CV Crop by BBoxes or CV Crop by Masks instead: they're hand-written, they take real bbox/mask data, and they'll do the boring integer version of this job without you hand-computing a centre.

Categoryimage/CV/low-level/cv2 G

Inputs (4)

NameTypeDefaultDescription
imageCOMFY_MATCHTYPE_V3Source image. The image output(s) echo this input's format. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
patchSizeCV_TUPLE0,0Size of the extracted patch. One value with 2 components (w, h) - it travels as a whole, so it cannot arrive half-connected. Wire it from 'CV Tuple' or type the components in place.
centerCV_TUPLE0,0Floating point coordinates of the center of the extracted rectangle within the source image. The center must be inside the image. One value with 2 components (x, y) - it travels as a whole, so it cannot arrive half-connected. Wire it from 'CV Tuple' or type the components in place.
patchTypeoptCOMBOsame as inputDepth of the extracted pixels. By default, they have the same depth as src .

Outputs (1)

NameTypeDescription
resultCOMFY_MATCHTYPE_V3Echoes the 'image' input's format: an IMAGE link comes back as IMAGE, MASK as MASK, NPARRAY stays NPARRAY.