Nodes/ComfyUI CV/CV MediaPipe Hand Pose
ComfyUI Node

CV MediaPipe Hand Pose

21 keypoints per hand, straight into your graph

By bmad4ever·Created 4 months ago·Updated 15 days ago· 1
CV MediaPipe Hand Pose
  • image
  • palms
  • bboxes
  • landmarks
  • world_landmarks
  • handedness
  • scores
  • hand_count
  • connections
  • keypoint_names
◄model▾►
◄conf_threshold0.90►

Stage two of MediaPipe's hand pipeline. Stage one (CV MediaPipe Palm Detect) gave you palms with an orientation; this node takes those crops and regresses the 21 hand keypoints per hand - wrist, thumb through pinky, four joints each. If you want a hand skeleton for ControlNet-style conditioning, for a geometric check on a generation, or as input to a hand-specific region mask, this is the bit that produces the numbers.

It runs through cv2.dnn on the ONNX export, which is the pack's whole premise and also its practical virtue here: Apache-2.0 all the way down, no ultralytics, no AGPL, no CUDA-specific build. Whether it's fast is another matter, and worth setting expectations on - this is CPU inference, one image at a time.

What the two stages actually buy you

You can't skip stage one. The palm output is what seeds each hand's crop and rotation, which is precisely what makes the landmark stage robust to hands being rotated or partly out of frame. Feed this node a bare image with no palms and there's nothing for it to do - the two nodes are a pair, not alternatives. This is the standard MediaPipe architecture, ported faithfully (the pack credits the OpenCV Zoo's handpose_estimation_mediapipe reference for exactly this code, Apache-2.0, with the models from Google).

Inputs and outputs

Required: image (the same image the palms were detected on - a batch uses its first frame, an NPARRAY counts as one frame), palms (the (N,19) array from CV MediaPipe Palm Detect), and model. Optional: conf_threshold (default 0.9 - palms whose confidence falls below this are dropped, not errored).

Outputs, and this node is generous with them:

  • bboxes - one {x, y, width, height, score} per hand, the tight hand box. Goes to core Draw BBoxes.
  • landmarks - (N·21, 2) float32 screen keypoints of every hand, in MediaPipe order (wrist, then thumb to pinky). This is the one you'll use most.
  • world_landmarks - (N·21, 3) metric 3D keypoints in metres, origin at the hand centre. This is the output people don't expect from a 2-D detector, and it's genuinely useful: real hand-sized measurements rather than image-space guesses.
  • handedness - (N,) float in [0,1]; ≤0.5 left, >0.5 right. Note it's a score, not a label, and the pack leaves the interpretation to you in the same order as bboxes.
  • scores, hand_count.
  • connections - (N·21, 2) int32 edge index-pairs of the MediaPipe hand skeleton, already offset per hand so it indexes landmarks directly. Feed landmarks + connections into CV Draw Connections to draw the bones.
  • keypoint_names - (N·21,) strings (wrist, thumb_tip, …), row-aligned with landmarks, for CV Draw Labels.

That trio - landmarks, connections, keypoint_names - is the nice part of the design. Drawing a skeleton in most stacks means knowing the MediaPipe edge list yourself; here it's an output.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv

Manager → ComfyUI CV, then the model:

ComfyUI/models/onnx/
  handpose_estimation_mediapipe_2023feb_int8bq.onnx

From OpenCV Zoo (Apache-2.0). The node's model combo lists everything under models/onnx, so it appears once copied.

Common issues

"The model is missing" on a workflow that ran yesterday. These ONNX files never ship with the pack and aren't in example_inputs/. They're recorded in the repo's model_sources.txt with download URLs and licences - that's the file to read when something opens red.

Empty outputs. Either no palms passed conf_threshold (0.9 is strict; drop it to 0.5 and see) or the palms came from a different image than the one you're passing in. Zero hands is a valid result - gate on hand_count rather than assuming.

Confusing the two coordinate sets. landmarks is in image pixels; world_landmarks is in metres with the origin at the hand centre. Overlaying requires projecting the latter, and forgetting which is which produces overlays that are off by a factor of several hundred.

Expecting a ControlNet-ready hand preprocessor. The classic OpenPose hand pipeline produces a rendering on a 512 canvas. This gives you data - points, connections, names. Wiring it into conditioning is your job, and if that's the goal, CV Draw Connections plus a resize is the shape of it. MediaPipe is the licence-clean stand-in for a lot of detection jobs in this ecosystem, but it supplies detection and landmarks, not an identity embedding - if you're doing face work, it can't replace ArcFace-based nodes like InstantID or PuLID, and it isn't trying to.

Categoryimage/CV/dnn

Inputs (4)

NameTypeDefaultDescription
imageNPARRAY,IMAGEThe SAME image the palms were detected on. An IMAGE batch uses its first frame; an NPARRAY is treated as one frame. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
palmsNPARRAYThe (N, 19) 'palms' array from 'CV MediaPipe Palm Detect'. Each row seeds one hand's crop/rotation.
modelCOMBOMediaPipe hand-pose .onnx model from ComfyUI/models/onnx (handpose_estimation_mediapipe_2023feb_int8bq.onnx).
conf_thresholdFLOAT0.900–1Minimum model confidence to keep a hand. Palms below this are dropped (not an error).

Outputs (8)

NameTypeDescription
bboxesBOUNDING_BOXOne {x, y, width, height, score} dict per hand (the tight hand box) - feed the core 'Draw BBoxes' node.
landmarksNPARRAY(N*21, 2) float32 screen keypoints of every hand, in MediaPipe order (wrist, thumb..pinky). Feed 'CV Draw Points'. Empty (0, 2) when no hands.
world_landmarksNPARRAY(N*21, 3) float32 metric 3D keypoints (x, y, z in meters, origin at the hand centre). For inspection / 3D use.
handednessNPARRAY(N,) float32 in [0, 1]: <=0.5 is a left hand, >0.5 right, same order as bboxes.
scoresNPARRAY(N,) float32 confidence per hand, same order as bboxes.
hand_countINTNumber of hands kept (palms that passed conf_threshold).
connectionsNPARRAY(N*21, 2) int32 edge index-pairs (the MediaPipe hand skeleton), already offset per hand to index the 'landmarks' output directly. Feed 'landmarks' + this into 'CV Draw Connections' to draw the bones. Empty (0, 2) when no hands.
keypoint_namesNPARRAY(N*21,) string array naming every keypoint (wrist, thumb_tip, ...), aligned row-for-row with 'landmarks'. Feed 'landmarks' + this into 'CV Draw Labels' to annotate them. Empty (0,) when no hands.