Nodes/ComfyUI-ZhiHui/🎯 智绘_提示词抠图(实验)
ComfyUI Node

🎯 智绘_提示词抠图(实验)

The prompt-based cutout that's still a stub — read this before wiring it up

By zhuyungen·Created 8 months ago·Updated 4 months ago· 0
🎯 智绘_提示词抠图(实验)
  • images
  • RGB图像
  • Alpha蒙版
  • 原图(透传)
◄promptperson►
◄confidence_threshold0.25►

Let me save you twenty minutes. ZH_BackgroundRemoverWithPrompt (🎯 智绘_提示词抠图(实验)) is supposed to be the thing where you type "person" or "the red car" and it cuts out exactly that object. The "(实验)" - experimental - label on the display name is doing real work here, because in the shipped code this node does not actually segment anything. Its segment_by_text function is a stub: feed it any image, any prompt, any threshold, and it returns the input image unchanged, a solid black (zero) mask, and the original image again. Nothing is looked at, nothing is matched, nothing is cut.

That's not me guessing from a thin README. It's in the source, right in nodes/image/background_remover.py - the entire implementation is a few lines that build an all-zeros mask tensor and pass your images straight through. I'm telling you because the node looks functional: it has a prompt box, a confidence_threshold slider, and three properly-named output sockets (RGB图像, Alpha蒙版, 原图透传). The shape is complete. The brain is missing.

What it should have been

Prompt-driven segmentation is a real, solved-in-principle pattern - it's GroundingDINO-class text-conditioned detection feeding a SAM-class segmenter, or a VLM that interprets "person on the left" and produces the mask. That's the "targeted masking" approach the KB's background-removal doc distinguishes from plain foreground/background removal, and it's genuinely the right tool when you want one object out of a scene rather than "everything that's in front." A workflow that does this properly in ComfyUI would look like: GroundingDINO node → box prompt → SAM node → mask out. This node was clearly aiming at that niche and hasn't shipped it.

The inputs, for completeness

In case a future version actually implements it: images in, prompt (default "person") as the text selector, and confidence_threshold (default 0.25) as the match cutoff. Outputs are RGB图像 (the cutout), Alpha蒙版 (the mask), and 原图透传 (input passed through untouched).

Install

It's part of the 智绘灵箱 (ComfyUI-ZhiHui) pack:

cd ComfyUI/custom_nodes
git clone https://github.com/zhuyungen/ComfyUI-ZhiHui.git

Restart ComfyUI (or use ComfyUI Manager, "智绘灵箱" / "ComfyUI-ZhiHui"). The node carries no extra dependencies of its own - it's pure pass-through - so install is trivial. Which is the only thing that's trivial about it right now.

What to do instead

Don't build a workflow on this node yet. If the author lands the real implementation in a later version, the API shape above is what you'll plug into. Until then, for prompt-targeted cutouts reach for the pack's own ZH_BackgroundRemover (background removal per model, not per prompt) or a dedicated GroundingDINO/SAM node pack - that's the combo that actually cuts "the red car" out of a street scene today.

This is a genuinely useful note to find ahead of time: on comfy.icu, the "(实验)" tag and 0 installs told you it was early, and the source confirms it's earlier than the UI suggests. Wire it up and you'll get a black mask and a perfectly untouched image - which, silently, is the worst kind of failure, because nothing errors out.

Category智绘灵箱/图片

Inputs (3)

NameTypeDefaultDescription
imagesIMAGE—
promptSTRINGperson—
confidence_thresholdFLOAT0.250–1—

Outputs (3)

NameTypeDescription
RGB图像IMAGE—
Alpha蒙版MASK—
原图(透传)IMAGE—