Nodes/comfyui-superside-nodes/Superside Florence-2 Smart Region Selector
ComfyUI Node

Superside Florence-2 Smart Region Selector

Ask Florence-2 to find the face, and get a mask plus crop coordinates back

By Superside·Created 3 months ago·Updated 6 days ago· 1
Superside Florence-2 Smart Region Selector
  • image
  • mask
  • mask_image
  • info
  • center_x
  • center_y
  • crop_width
  • crop_height
◄region_typeface►
◄api_key►
◄custom_text►
◄selection_modelargest►
◄padding_percent8.0►
◄return_rect_maskfalse►
◄mask_blur_percent0.0►
◄detection_modeauto►
◄upload_max_dimension2048►

This is the automation half of the region-detail loop: instead of hand-painting a mask for "just the face," you tell this node what you want and it finds it, segments it, and hands you a mask plus the exact crop coordinates. It runs Florence-2's grounding/segmentation on fal, and it's built to feed the pack's Crop By Region + Stitch Region pair - you select the face here, crop just the face, render it at a resolution the model can actually resolve, and paste it back. The KB calls this kind of thing the "All detailers are doing is zooming in on a mask" loop, and this is the detector step of it.

It selects one semantic region at a time from a fixed set - face, upper_body, lower_body, full_body, or object. For object you supply custom_text ("red sneaker", "logo on the sleeve") and Florence-2 grounds that phrase to a region. selection_mode picks whether you want the largest instance (default, when there are several people) or to merge_all of them into one mask. padding_percent (default 8) pads the region out so your downstream crop has context, and return_rect_mask swaps the soft segmentation mask for a hard rectangle if that's what the next step wants.

The outputs are the important part, because they're the shared contract that makes this node interchangeable with the pack's SAM 3 selector:

  • mask (MASK) - the segmented region, ready for masking/inpainting.
  • mask_image (IMAGE) - the same mask as an image, if you need to inspect it or convert.
  • info (STRING) - JSON of what was selected.
  • center_x, center_y, crop_width, crop_height (INT) - the region's center and size in original-image pixels. These feed Crop By Region's four matching inputs directly.

That last quartet is what makes the whole pipeline coordinate-exact: the crop node returns the precise crop rectangle it used, and the stitch node pastes the processed result back at the same spot in the full-res original. No guessing, no resizing the original.

The model itself is the same Florence-2 the caption node uses, but here it's doing grounding - matching text to a location - which is its genuinely strong task, not its weak one. One honest tradeoff to carry in: this is a single-region selector with a fixed vocabulary of five options. If you need "glasses frame without the lens" or "the third car's headlight," that's the pack's SAM 3 node, which has 19 presets and takes a box prompt to focus a sub-part. Florence's vocabulary is the five built-ins plus whatever you type into custom_text, and for the body-region basics it's fast and reliable.

It's an API call (fal credits, image leaves your machine - paste your key into api_key, blank falls back to FAL_KEY). The alternative is a local selector: the KB documents plenty of open detectors, but those need their own installs; this one is zero-setup and sits right next to the rest of the pack.

Install - ComfyUI Manager (search "comfyui-superside-nodes") or:

cd ComfyUI/custom_nodes
git clone https://github.com/Superside/comfyui-superside-nodes
pip install -r requirements.txt

Restart ComfyUI, look under Superside. No model downloads.

CategorySuperside

Inputs (10)

NameTypeDefaultDescription
imageIMAGE—
region_typeCOMBOface6 options: glasses, face, upper_body, lower_body, full_body, object
api_keySTRING—
custom_textoptSTRING—
selection_modeoptCOMBOlargest2 options: largest, merge_all
padding_percentoptFLOAT8.00–100—
return_rect_maskoptBOOLEANfalse—
mask_blur_percentoptFLOAT0.00–100Soft blur applied to the returned mask, as a percentage of the region's longest edge. Florence's segmentation of a thin object like a spectacle frame is jagged; a few percent gives a downstream stitch or inpaint something it can blend against.
detection_modeoptCOMBOauto'auto' tries segmentation and falls back to a grounding box, except for a face, where the box is tried first because segmentation tends to grab the whole head. 'segmentation' and 'grounding_bbox' force one or the other.
upload_max_dimensionoptINT2048512–4096Downscale the image's longest edge before uploading it for detection. Lower it if fal closes the connection on big images.

Outputs (7)

NameTypeDescription
maskMASK—
mask_imageIMAGE—
infoSTRING—
center_xINT—
center_yINT—
crop_widthINT—
crop_heightINT—