FeiHou Easy H3 Face Refine (Experimental)
A second sampling pass, not a sharpening filter
- images
- model
- h3_context
- reference_face
- reference_faces
- audio
- images
- audio
- fps
- report
The name says "refine", which sounds like a filter. It isn't. It crops the face out of every frame, runs H3 itself over those crops again at a fresh denoise, then blends them back into the video. The author's own one-liner is the honest summary: experimental H3 face crop/refine/stitch. Adds sampling; does not guarantee lip sync. Original audio passes through.
Why you'd want it
A face occupying 70 pixels of a frame comes out a smear no matter how good the model is - the latent has no budget to spend there. The fix has always been to give that region its own pass: the detect-crop-re-render-paste loop behind ADetailer and FaceDetailer. It matters more in video, where distant faces mangle as the shot moves and your existing option is a per-frame image detailer fighting temporal consistency.
So this is ADetailer's loop, with H3 doing the refinement instead of a still-image checkpoint - the only sensible choice, since the refined crop has to agree with its neighbours.
How it works, mechanically
You hand it the finished frames plus the same H3 model and the h3_context the main node produced. From there:
- Detect. A YOLO face detector runs per frame, boxes are smoothed over 21 frames, and shot cuts are found with PySceneDetect so the tracker doesn't glide a box across a hard cut - it fails loudly rather than smoothing through one.
- Pick whose face.
largest_facetakes the biggest detection;reference_identitytracks against a reference face instead, which is the mode that needs InsightFace embeddings. - Crop and pad. Crops go to a 512 or 768 canvas (edge-extended, not stretched, so nothing distorts), then get padded onto H3's
17k+5frame grid. - Re-sample. The crop runs through the same reference-conditioning path the main node uses - denoise 0.25, 8 steps, BasicScheduler guider - then gets VAE-decoded.
- Stitch. Refined frames are averaged back in with 17-frame overlap between chunks, then pasted through a feathered face-ellipse mask with colour matching. Where the detector lost the face, the original fades back in.
Output frame count, size and FPS are unchanged - it hard-errors if they drift. And the original audio always passes through untouched: audio_lock feeds aligned audio into the sampling as conditioning (H3's native audio mask, video mask 1, audio mask 0). It isn't a lip-sync model, and your output audio is the audio you plugged in.
Inputs and outputs worth caring about
Required: images (the final decoded frames - plug this in before any frame interpolation), model (the H3 model doing the refining; LoRAs load externally, no Turbo is added), and h3_context from the Easy H3 main node or loader. Optional: reference_face and audio.
reference_face falls back to the first reference image the main node kept in its context, so if you generated with one you can leave it unconnected.
Outputs are images, audio, fps and report. The first three go to your video save node; report goes to a text preview, and you should actually read it - it says how many segments were sampled and what the reference-face logic decided.
Two widget-panel quirks. Most of the interesting controls (canvas_size, confidence, crop_factor, chunk_frames, sampler, scheduler, audio_lock, clean_second_model, force_offload) only appear with advanced on. And the prompt box is hidden permanently - the default text is still what gets sampled, and anything you saved there still counts. You can't see it, but it's running.
Two switches earn their own paragraph. clean_second_model (default on) swaps to the pre-LoRA base the loader recorded, so the refine pass doesn't sample through your LoRA stack and external patches; it needs that recorded attachment, and with an external model chain it tells you to re-run the updated loader/main node or turn it off. chunk_frames offers 不分段, 240, 192, 120 and 72 - the label is Chinese for "don't split" and it's the default. It only removes length-based splitting, so long takes can still be a VRAM peak.
Installing it
The node ships in the pack; its dependencies are separate and optional.
cd ComfyUI/custom_nodes
git clone https://github.com/FX-FeiHou/ComfyUI-FeiHou-Easy-H3
# use the SAME python that runs ComfyUI:
python -m pip install -r ComfyUI/custom_nodes/ComfyUI-FeiHou-Easy-H3/requirements-face-refine.txt
That pulls ultralytics, scipy and scenedetect>=0.7. Drop detector weights (face_yolov8m.pt and friends) in ComfyUI/models/ultralytics/bbox/; the node also picks up detectors registered by Impact Pack's ultralytics_bbox folder. reference_identity needs InsightFace plus an ONNX Runtime and recognition weights - one provider, CPU or GPU, never both, and check those weights' licensing yourself. Missing deps only break this node; the rest of the pack still loads, with a message pointing at requirements-face-refine.txt.
Where it bites
It's a prototype. The author's own boundary list: single person, segmented by shot cuts, chunk seams still need checking, and this version is a CPU/code regression prototype with no full GPU quality validation.
24 FPS only. It raises if the incoming FPS isn't 24, so put it before interpolation, not after.
It can change the face. Even at denoise 0.25 with audio_lock off, identity, expression, features and mouth shape can move. Keep the original take - enabled off or denoise at 0 is a passthrough, so A/B is cheap.
No face found is not an error. You get your original frames back plus a line in report. Before assuming it's broken, check confidence (0.35) and crop_factor (2.5).
audio_lock needs properly aligned audio. The final video's own audio, cut from frame 0 - not a whole song, not the unaligned reference track. It errors on the wrong shape, and handing it a full-length track is an easy mistake.
VRAM is still your problem. force_offload (default on) pushes detectors and the refine model/VAE out of VRAM afterwards, but the author promises nothing at every resolution. Drop canvas_size to 512, or chunk, when it OOMs.
Inputs (25)
| Name | Type | Default | Description |
|---|---|---|---|
| images | IMAGE | — | |
| model | MODEL | — | |
| h3_context | MINIMAX_H3_CONTEXT | — | |
| enabled | BOOLEAN | true | — |
| detector | COMBO | 1 options: face_yolov8m.pt | |
| target | COMBO | 2 options: largest_face, reference_identity | |
| denoise | FLOAT | 0.250–1 | — |
| steps | INT | 81–100 | — |
| seed | INT | 00–18446744073709550000 | — |
| advanced | BOOLEAN | false | — |
| canvas_size | COMBO | 768 | 2 options: 512, 768 |
| confidence | FLOAT | 0.350.05–0.95 | — |
| crop_factor | FLOAT | 2.51.2–5 | — |
| chunk_frames | COMBO | 不分段 | 5 options: 不分段, 240, 192, 120, 72 |
| sampler_name | COMBO | euler | 44 options: euler, euler_cfg_pp, euler_ancestral, euler_ancestral_cfg_pp, heun, heunpp2, +38 |
| scheduler | COMBO | simple | 9 options: simple, sgm_uniform, karras, exponential, ddim_uniform, beta, +3 |
| prompt | STRING | Refine the face in this video crop. Preserve identity, expression, gaze, mouth movement, lighting and motion. Natural facial detail; no new motion or camera changes. | — |
| force_offload | BOOLEAN | true | — |
| audio_lock | BOOLEAN | false | Experimental: condition face sampling on the connected, time-aligned original audio. Not a guaranteed lip-sync correction. |
| clean_second_model | BOOLEAN | true | Use the recorded pre-LoRA second-pass base. Requires an Easy H3 second-pass model; removes external patches as well. |
| face_mode | COMBO | 单人 | 单人:修复一条人脸轨迹。多人:逐脸采样并按每张脸自己的遮罩合成。 |
| max_faces | INT | 22–8 | Maximum faces in multi mode. Sampling is sequential to keep peak VRAM stable. Without a reference batch, ranked faces are tracked. |
| reference_faceopt | IMAGE | — | |
| reference_facesopt | IMAGE | — | |
| audioopt | AUDIO | — |
Outputs (4)
| Name | Type | Description |
|---|---|---|
| images | IMAGE | — |
| audio | AUDIO | — |
| fps | FLOAT | — |
| report | STRING | — |