cv2.grabCut
The classic CV answer to \u201ccut this thing out\u201d
- img
- mask
- bgdModel
- fgdModel
- mask
- bgdModel
- fgdModel
GrabCut is the 2004 answer to "give me a rough box around the object and I'll find the real edge." You draw a rectangle - or scribble a bit - and it iterates a Gaussian mixture colour model, deciding pixel by pixel what is likely foreground, tightening the boundary along the way. It is deterministic, CPU-only, needs no model download, and it is terrible at exactly what modern matting models are great at: hair, veils, glass, motion blur, anything semi-transparent.
So treat this node as a region tool, not a cutout tool. It is excellent at turning a rough box into a tight polygon-ish mask you then hand to inpainting, to a crop policy, or to a compositing step. If you want a genuine alpha matte for a person against a busy background, a learned matting model is a different output type entirely - segmentation gives you foreground/background, matting gives you fractional alpha, and asking the wrong one for a veil is a structural mismatch, not a quality problem (background-removal.md).
How it works
cv2.grabCut(img, mask, rect, bgdModel, fgdModel, iterCount, mode). The mask is the state: each pixel is one of sure background, sure foreground, probable background, probable foreground. In GC_INIT_WITH_RECT mode OpenCV initialises that mask itself from your rectangle - everything outside becomes sure background, everything inside starts as probable foreground - then runs a few iterations of colour-model fitting and min-cut segmentation. Further calls with GC_INIT_WITH_MASK or GC_EVAL refine an existing mask using it as the starting point, which is where the model arrays come in.
Those models are the reason this is not a fire-and-forget node. bgdModel and fgdModel are data buffers - the author's tooltip is blunt about it: a data array, not an image, so only an NPARRAY link is accepted. Wire the outputs of one call into the inputs of the next to keep refining, and do not hand-edit them.
Because grabCut is genuinely slow, the pack runs it in a separate interruptible subprocess, so ComfyUI can cancel mid-run when you realise you pointed it at 4K frames.
Inputs and outputs
img- 8-bit 3-channel. Colour wins here; grayscale images give the GMM very little to separate on.mask- 8-bit single channel, the state array. ForGC_INIT_WITH_RECTthe function initialises it, so you mostly just supply something sized like the image.rect_x,rect_y,rect_w,rect_h- the initial box, split into four widgets. Everything outside it is declared sure background, so a rectangle that clips the object breaks the result immediately.bgdModel,fgdModel- the model buffers described above. Required, NPARRAY-only.iterCount- 1 is enough to see the idea, 5 is the usual value, more than 10 rarely changes anything.mode- dropdown of the four GrabCut modes:GC_INIT_WITH_RECT(default),GC_INIT_WITH_MASK(respect your scribbles),GC_EVAL,GC_EVAL_FREEZE_MODEL.
Outputs are mask (the raw label map, values 0–3, view it with Preview CV Array in heatmap mode), plus bgdModel and fgdModel.
The curated node is almost certainly what you want
The pack also ships CV GrabCut, and it does the translation work this raw wrapper leaves to you: it takes 2+ points to derive the rectangle, accepts a mask as probable foreground, accepts separate sure_foreground / sure_background scribble masks, runs the iterations, and returns a ready 0/255 MASK plus the label map and a success flag that comes back false instead of erroring when your hints were unusable. It exists precisely because the raw signature is fiddly - that label map is not a mask, and iterCount, rectangle and mode all have to agree.
Use the raw node when you are chaining iterations by hand or porting an OpenCV sample; use CV GrabCut when you just want the subject.
Install
Manager → search comfyui_cv (bmad4ever), or:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"
Python ≥ 3.12 with a recent ComfyUI on the V3 node API. No models. Just do not let another package replace your contrib wheel - tools/repair_opencv_contrib.py --check / --apply is for when one did.
When it goes wrong
- Everything becomes background. Classic symptom of a rectangle that does not contain the object, or an object whose colour is close to the background. GrabCut has no semantic knowledge; it only has your box and colour statistics.
- Fuzzy, blobby boundary. Too few iterations, or the object's boundary is low-contrast. Add scribbles via the curated node instead of an empty mask - a few strokes of sure-foreground change the result more than more iterations.
- Two hundred milliseconds became twenty seconds. grabCut is iterative min-cut; cost scales with area, not just iterations. Preview on a downscaled copy, then run the final pass.
- A hard mask where you wanted an alpha edge. That is the tool working as designed. Feed the mask into inpainting or compositing, or step up to a matting model (
masking-detection-detailing.md).
Inputs (10)
| Name | Type | Default | Description |
|---|---|---|---|
| img | NPARRAY,IMAGE,MASK | Input 8-bit 3-channel image. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| mask | NPARRAY,IMAGE,MASK | Input/output 8-bit single-channel mask. The mask is initialized by the function when mode is set to #GC_INIT_WITH_RECT. Its elements may have one of the #GrabCutClasses. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| rect_x | INT | 0-2147483648–2147483647 | Rectangle top-left corner X in pixels. |
| rect_y | INT | 0-2147483648–2147483647 | Rectangle top-left corner Y in pixels. |
| rect_w | INT | 00–2147483647 | Rectangle width in pixels (>= 0). |
| rect_h | INT | 00–2147483647 | Rectangle height in pixels (>= 0). |
| bgdModel | NPARRAY | Temporary array for the background model. Do not modify it while you are processing the same image. A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here. | |
| fgdModel | NPARRAY | Temporary arrays for the foreground model. Do not modify it while you are processing the same image. A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here. | |
| iterCount | INT | 0-2147483648–2147483647 | Number of iterations the algorithm should make before returning the result. Note that the result can be refined with further calls with mode==#GC_INIT_WITH_MASK or mode==GC_EVAL . |
| modeopt | COMBO | GC_INIT_WITH_RECT | Operation mode that could be one of the #GrabCutModes |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| mask | NPARRAY | — |
| bgdModel | NPARRAY | — |
| fgdModel | NPARRAY | — |