Nodes/face_mosaic/Torch(MTCNN) 视频人脸马赛克
ComfyUI Node

Torch(MTCNN) 视频人脸马赛克

MTCNN video face mosaic — the lightweight middle ground

By yzzky·Created 12 months ago·Updated 12 months ago· 0
Torch(MTCNN) 视频人脸马赛克
    • output_video_path
    • processing_info
    ◄video_path►
    ◄mosaic_size20►
    ◄mosaic_type▾►
    ◄output_format▾►
    ◄use_gputrue►
    ◄output_path►

    Between the pack's ancient Haar detector and its heavyweight InsightFace option sits a nice middle: MTCNN, run through PyTorch. MTCNN is a three-stage CNN cascade that's been the default academic face detector for years - better on angles and lighting than Haar, way lighter on your system than the full InsightFace stack. If you want better-than-2001 detection without pulling in the whole face-swap ecosystem, this is the node. In the menu it's Torch(MTCNN) 视频人脸马赛克 under YZZ_Face_Mosaic/Torch.

    How it works

    The node builds an MTCNN instance from the facenet_pytorch package with keep_all=True - meaning it returns every face it finds, not just one - and runs it on CUDA if use_gpu is on and a device exists, CPU otherwise. Each video frame becomes a PIL image, the detector returns bounding boxes, and the same pixelate / blur / black_box painter used across this pack fills them in. Output goes to ComfyUI/output/face_mosaic_torch/ as output_video_path, with processing_info giving you frames processed, faces found, and real processing fps.

    MTCNN is notably light: the model files are tiny compared to InsightFace's buffalo_l, so first-run setup is quick and the memory footprint stays low.

    Inputs

    • video_path - a string path; wire it from the pack's VideoUploadNode rather than typing.
    • use_gpu - on by default; it only actually uses CUDA when available.
    • mosaic_size and mosaic_type - the usual three styles.
    • output_format (mp4 / avi / mov) and optional output_path.

    Outputs: output_video_path (STRING) and processing_info (STRING JSON).

    Installing it - the one extra step

    MTCNN comes from facenet-pytorch, which is not in the pack's requirements.txt. Install the pack, hit this node, and you'll see a "facenet-pytorch 未安装" warning in the console - and the node will detect no faces, writing out a clean un-mosaicked copy. That's the trap. Fix it with:

    pip install facenet-pytorch
    

    in your ComfyUI Python environment, then restart. The pack itself installs via ComfyUI Manager (yzz_face_mosaic) or git clone https://github.com/yzzky/yzz_face_mosaic into custom_nodes, then pip install -r requirements.txt.

    Where it bites

    MTCNN is better than Haar, not magic - extreme angles, very small faces, and heavy occlusion will still slip through, and it's slower per frame than the CPU Haar node. There's also a known quirk where occasional frames return None instead of boxes; the code swallows those and moves on, so you may see the odd flicker. For the record, the pack also has a Torch image twin (TorchImageFaceMosaicNode) with the same detector if you only need stills. If you want the absolute best detection and don't mind a bigger download, the InsightFace-based GPU nodes edge it out - but for most masking work, this is the sweet spot.

    CategoryYZZ_Face_Mosaic/Torch

    Inputs (6)

    NameTypeDefaultDescription
    video_pathSTRING—
    mosaic_sizeINT205–100—
    mosaic_typeCOMBO3 options: pixelate, blur, black_box
    output_formatCOMBO3 options: mp4, avi, mov
    use_gpuBOOLEANtrue—
    output_pathoptSTRING—

    Outputs (2)

    NameTypeDescription
    output_video_pathSTRING—
    processing_infoSTRING—