Nodes/ComfyUI CV/CV YuNet Face Detect
ComfyUI Node

CV YuNet Face Detect

Faces and Five Landmarks, Minus the InsightFace Licence

By bmad4ever·Created 4 months ago·Updated 15 days ago· 1
CV YuNet Face Detect
  • image
  • bboxes
  • landmarks
  • scores
  • face_count
◄model▾►
◄conf_threshold0.90►
◄nms_threshold0.30►
◄top_k5000►

CV YuNet Face Detect finds faces and, more usefully, hands back five landmarks per face. It's a small, fast CNN - "a tiny millisecond-level face detector" is the paper's own framing - running through OpenCV's FaceDetectorYN, inside a ComfyUI node.

The licence is half the reason to care

If you only need to find and crop a face, InsightFace is the wrong tool - MIT code, non-commercial weights, and that split quietly makes IP-Adapter FaceID, InstantID and PuLID unsellable. Detection is the easy escape route.

YuNet is exactly that swap. The model file is MIT (OpenCV Zoo's packaging of Shiqi Yu's libfacedetection training work), and the companion embedding node in this same pack uses SFace, which is Apache-2.0. So YuNet → SFace → CV Embedding Match is a complete detect-and-recognise pipeline with licences you can actually ship, and the pack's own workflow 66 does precisely that.

What YuNet can't do is tell two faces apart. It has no identity embedding, just boxes, points and confidences.

How it works

This is one of the pack's curated nodes, not a raw cv2.* wrapper - the generator that produces the other ~470 nodes only parses top-level OpenCV functions, so a class-based detector like cv2.FaceDetectorYN needs a hand-written bridge.

The node builds a detector with your confidence, NMS and top_k settings, sets the input size to each frame's own size, and runs detection. Each returned row carries a box, ten floats (the five landmarks as x/y pairs), and a confidence. The node unpacks that into one bbox dict per face, an (N, 2) landmark array, (N,) scores and a total count. Landmark order is right eye, left eye, nose tip, right mouth corner, left mouth corner. Coordinates are in the image you fed - no letterbox maths, nothing to un-scale.

Those landmarks aren't decoration. cv2.alignCrop reads exactly those ten floats out of a YuNet row, which is how the pack's SFace node produces its aligned 112×112 crops. Detection and alignment, one data type, no glue.

Inputs and outputs

  • image - an IMAGE batch is processed frame by frame; an NPARRAY is treated as one frame.
  • model - picked from a dropdown listing the .onnx files under ComfyUI/models/onnx.
  • conf_threshold - defaults to 0.9, which is high. Lower it (0.6–0.8) if you're missing faces; you'll also collect more false positives.
  • nms_threshold (0.3) merges overlapping boxes; top_k (5000) caps how many candidates survive to NMS. Leave top_k alone.
  • bboxes - one {x, y, width, height, score} dict per face, which is the core BOUNDING_BOX type, so the stock Draw BBoxes and Image Crop nodes take it. A batch nests one list per frame.
  • landmarks - (N, 2) float32, feed it to CV Draw Points.
  • scores - per-face confidence in the same order as the boxes.
  • face_count - total across the batch, handy for a switch or a filename tag.

Zero faces is a legitimate result, not an error: you get an empty box list and a (0, 2) landmark array.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"

Manager users: search comfyui_cv. Python ≥ 3.12 and a ComfyUI new enough for the V3 node API. Keep the contrib wheel - a plain opencv-python install over it empties the contrib submodules.

Then grab face_detection_yunet_2023mar.onnx from opencv/face_detection_yunet on HuggingFace into ComfyUI/models/onnx. The dropdown is built when the page fetches node definitions, so drop the file in and reload the ComfyUI tab before wondering why it isn't listed.

Where people get burned

The model not appearing in the dropdown is nearly always one of two things: the file is not in models/onnx, or you didn't reload the page.

Then there's the cost model, which surprises people. DNN work here runs in a throwaway worker subprocess per node execution, so the model is created fresh on every queue run - a batch is where the loading gets amortised. If you're detecting across a folder of stills, feed them as one IMAGE batch rather than running the node fifty times.

Two smaller things. On a batch, bboxes nests one list per frame while landmarks and scores are flattened across frames, so if you need per-frame grouping, count the boxes per frame yourself. And this is a detector with five coarse points - no pose, no blendshapes.

Finally, mind the pack's own framing: a personal project written with heavy LLM assistance, updates not planned, not production software. The detector inside it is solid; the surrounding pack is a workbench.

Categoryimage/CV/dnn

Inputs (5)

NameTypeDefaultDescription
imageNPARRAY,IMAGEInput image. An IMAGE batch is processed frame by frame; an NPARRAY is treated as a single frame (gray/BGR/BGRA, any dtype). The landmark and score outputs are concatenated across frames. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
modelCOMBOYuNet .onnx model file from ComfyUI/models/onnx.
conf_thresholdFLOAT0.900–1Minimum confidence to keep a face. Lower detects more (and more false positives).
nms_thresholdFLOAT0.300–1Non-maximum suppression IoU: boxes overlapping by more than this are merged.
top_kINT50001–20000Keep at most this many boxes before NMS.

Outputs (4)

NameTypeDescription
bboxesBOUNDING_BOXOne {x, y, width, height, score} dict per face - feed the core 'Draw BBoxes' node or 'Image Crop'. A batch nests one list per frame.
landmarksNPARRAY(N, 2) float32 point array: the 5 landmarks of every face, in order right eye, left eye, nose, right/left mouth corner. Feed 'CV Draw Points'. Empty (0, 2) when no faces.
scoresNPARRAY(N,) float32 confidence per face, same order as bboxes.
face_countINTTotal number of faces detected across the batch.