Nodes/ComfyUI CV/CV MediaPipe Palm Detect
ComfyUI Node

CV MediaPipe Palm Detect

Hand detection that runs through OpenCV's DNN module

By bmad4ever·Created 3 months ago·Updated 14 days ago· 1
CV MediaPipe Palm Detect
  • image
  • bboxes
  • palms
  • palm_points
  • scores
  • palm_count
◄model▾►
◄score_threshold0.60►
◄nms_threshold0.30►
◄top_k5000►

Hands are the part of a person that generative models break first and most visibly, and the reason automatic detailing works on faces but often not on hands is that the detector is the hard half. There's a well-worn YOLO for faces. Hands are messier: fingers occlude each other, palms rotate freely, and a bounding box around a hand isn't a thing you can crop and re-render usefully without knowing its orientation.

MediaPipe's hand pipeline handles that with a two-stage design, and this node is stage one. It's a palm detector - not a hand detector - and that's deliberate, as you'll see.

Why palm, not hand

A palm is rigid and roughly planar, so a rotated rectangle around a palm gives you a crop and an orientation you can normalize a hand into. Fingers are articulated, so a "hand" box would be a badly-defined container. Detecting the palm first and then regressing landmarks inside it is what makes the landmark stage work on rotated hands, occluded fingers and both hands at once - up to a couple of dozen of them in MediaPipe's own formulation.

Second reason this node exists: it runs entirely through cv2.dnn, on an ONNX export of MediaPipe's palm SSD. No ultralytics, no AGPL, no PyTorch model graph. MediaPipe is Apache-2.0, which is exactly why it shows up in this ecosystem whenever someone needs a detection licence they can live with. If you're building something you intend to sell, that's the whole argument.

Third: the anchors and the sigmoid decoding are implemented here. The raw ONNX gives you a blob of box/score predictions on a grid; turning that into boxes needs NMS and anchor decoding, which is why score_threshold, nms_threshold and top_k are inputs.

Inputs and outputs

Required: image (an NPARRAY or a ComfyUI IMAGE - a batch uses its first frame; an NPARRAY in gray/BGR/BGRA, any dtype, is treated as one frame) and model.

Optional: score_threshold (default 0.6 - lower detects more, including more false positives), nms_threshold (default 0.3 - palms overlapping by more than this IoU get merged), top_k (default 5000, the candidate cap before NMS).

Outputs:

  • bboxes - a BOUNDING_BOX value with one {x, y, width, height, score} per palm, which goes straight into core Draw BBoxes.
  • palms - (N, 19) float32: each row is [x1, y1, x2, y2, 7×(lx, ly), score]. This is what you wire into CV MediaPipe Hand Pose. It's not a general-purpose array; it's the handoff format between the two stages.
  • palm_points - (N·7, 2) the seven palm landmarks (wrist plus finger bases) of every palm, for CV Draw Points.
  • scores - (N,) confidence per palm.
  • palm_count.

Outputs are data only. This pack deliberately doesn't draw anything for you - no overlay image. Boxes go to Draw BBoxes, points go to CV Draw Points. If you were expecting an annotated image like a ControlNet preprocessor returns, that's the difference.

Zero palms is a valid result with empty outputs, never an error - the discipline across this whole pack.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv

Manager → ComfyUI CV. Then the model, which does not ship with the pack:

ComfyUI/models/onnx/
  palm_detection_mediapipe_2023feb.onnx

From the OpenCV Zoo (Apache-2.0). The _int8bq variant is also listed in the pack's model_sources.txt if you want the smaller one. The model COMBO lists every .onnx under models/onnx, subdirectories included, so it appears in the dropdown once the file is there - no restart needed for the file itself, but do refresh the page.

Pack dependencies: opencv-contrib-python-headless~=5.0.0.93, Python ≥ 3.12, V3-API ComfyUI. Nothing heavy.

Common issues

Red node / empty dropdown. The model file isn't where the node is looking. It's ComfyUI/models/onnx, not the pack folder.

Two hands detected, one is a face or a knee. Lower the score_threshold and this is what you get. 0.6 is a sane default precisely because the detector is over-eager at 0.3.

Only the first frame gets processed. Stated in the node's own description, and it catches people who feed a video batch and expect per-frame output. Loop or batch externally.

OpenCV wheel collisions. All four OpenCV distributions share one site-packages/cv2, so installing a plain opencv-python over the contrib wheel can empty contrib submodules and make contrib nodes disappear with no error. python tools/repair_opencv_contrib.py --check diagnoses it. If import cv2 itself fails with a missing shared library (libxcb.so.1 on minimal Linux/NixOS setups is the reported case), that's a system-library problem and the headless wheel is the usual fix - which is what this pack already depends on.

Categoryimage/CV/dnn

Inputs (5)

NameTypeDefaultDescription
imageNPARRAY,IMAGEInput image. An IMAGE batch uses its first frame; an NPARRAY (gray/BGR/BGRA, any dtype) is treated as one frame. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
modelCOMBOMediaPipe palm .onnx model from ComfyUI/models/onnx (palm_detection_mediapipe_2023feb.onnx).
score_thresholdFLOAT0.600–1Minimum sigmoid confidence to keep a palm. Lower detects more (and more false positives).
nms_thresholdFLOAT0.300–1Non-maximum suppression IoU: palms overlapping by more than this are merged.
top_kINT50001–20000Keep at most this many candidate boxes before NMS.

Outputs (5)

NameTypeDescription
bboxesBOUNDING_BOXOne {x, y, width, height, score} dict per palm - feed the core 'Draw BBoxes' node.
palmsNPARRAY(N, 19) float32: each row [x1, y1, x2, y2, 7*(lx, ly), score]. Feed this straight into 'CV MediaPipe Hand Pose'. Empty (0, 19) when no palms.
palm_pointsNPARRAY(N*7, 2) float32: the 7 palm landmarks of every palm (wrist + finger bases), for 'CV Draw Points'. Empty (0, 2) when no palms.
scoresNPARRAY(N,) float32 confidence per palm, same order as bboxes.
palm_countINTNumber of palms detected.