MiniMax H3 One-click Face Repair - Star7
Your H3 clip is fine except the face — this node fixes just the face
- sampled_av_latent
- refine_context
- refined_av_latent
- report
You spent twenty minutes on a clip and the shot is great - motion, framing, the audio syncs - and the face looks like it was painted with a wet brush. In image land you'd reach for an ADetailer-style pass: detect, crop, re-diffuse the crop, paste it back. Video gives you no such free lunch - what you have after sampling is a packed H3 latent holding video and audio as one nested tensor.
MiniMaxH3FaceRefineLatentStar7 (MiniMax H3 One-click Face Repair - Star7) is that detailer pass, built for H3's latent layout.
Why it needs a second cable
A second sampling pass needs your model chain, your positive conditioning and the video VAE - without you re-wiring the loader, the LoRAs and the chunk node a second time. That's what refine_context is: a STAR7_H3_REFINE_CONTEXT bundle emitted by MiniMax H3 All-in-one Conditioning - Star7 (same pack). Feed it anything else and it stops with Star7 face-repair context is incomplete. So the swap is: all-in-one conditioning instead of your usual one, plus a single wire.
All-in-one Conditioning "refine_context" --> Face Repair "refine_context"
Sampler "sampled_av_latent" --------------> Face Repair "sampled_av_latent"
Face Repair "refined_av_latent" ----------> Optional HD Upscale --> Chunked Decode
How the pass actually works
The node decodes the video member of the packed latent to frames, runs YOLO face detection with per-frame tracking (with automatic scene-cut detection, so tracks don't smear across cuts), and crops each track onto a square canvas - auto caps it at 768, or you can pin 512/768. It then builds a fresh H3 video latent at canvas size, keeps the original audio latent, applies a per-frame denoise mask, and samples just that crop: simple schedule, res_multistep, refine_steps steps.
The part worth understanding, because it changes how you wire the end of the graph: the refined crops are not decoded and re-encoded into the full video. The node packs each repaired crop into a compact overlay carrying the tracking transform and your blend/feather settings, attaches it to the latent, and hands back your original sampled_av_latent untouched - base video latent and audio preserved exactly. The RGB stitch happens once, later, in MiniMax H3 Chunked Decode.
A warning as much as a design note: a stock VAEDecode at the end of your graph ignores the overlay and hands you the original, un-repaired footage.
Inputs you'll actually touch
sampled_av_latent- the sampler's output, and it has to be the real thing. A plain image latent gets yourequires a sampled MiniMax H3 packed audio-video LATENT.refine_context- from the all-in-one conditioning node, as above.preset- the four named presets are tuned, not decorative; they set how hard the crop gets redrawn.自动平衡starts at 0.30 denoise,真人保真is gentler (0.25, tighter crop, 512 canvas),远景小脸is the wide-shot preset (0.48 denoise, 768 canvas),动漫角色sits at 0.32. Thecustom_*fields only apply when preset is自定义- the most common "why did nothing change" trap here.face_count(1–4) andtarget_face-主人物tracks the largest face,画面中央the centre-most, and参考图匹配needs insightface (otherwise it warns and falls back to main-subject tracking).refine_steps(1–12),seed,preserve_repair_detail- the last one defaults on and lifts video below about 1 MP to a 0.98–1.01 MP canvas before pasting, so the repaired crop isn't squashed on the way back in. It never downscales.
face_lora, face_lora_strength and face_attention all default to inheriting the first pass. Point face_attention at another backend and it applies to the repair pass only, then gets restored - handy when you want the faces running under a different backend than the bulk video.
Outputs and installing
Two outputs: refined_av_latent (LATENT) and report (STRING). Read the report: final resolution, the sigma the pass started from, the detector, the LoRA and attention in play - the fastest way to tell "it ran but the effect is subtle" apart from "it didn't run".
Install via Manager (search MiniMax H3 Activation Chunk - Star7), comfy node install minimax-h3-chunk-star7, or:
cd ComfyUI/custom_nodes
git clone https://github.com/star7code/minimax-h3-chunk-star7.git
The pack pulls in ultralytics (the detector runtime) and scenedetect. The face_detector dropdown defaults to face_yolov8m.pt; on first use the node downloads it from Bingsu/adetailer - HF mirror first, then HF - verifies the SHA-256 and drops it in ComfyUI/models/ultralytics/bbox. The same dropdown lists any other face detector you already have in that folder.
Where people get burned
It no-ops instead of erroring when it can't find a face. no face detected and no stable face track both log a warning and pass your latent straight through, so the report string is where you find out. For a wide shot with a genuinely tiny face, try 远景小脸 before deciding the node is broken.
Budget for it. It decodes the source video up front and runs one extra sampling pass per face - face_count=4 is four passes plus detection and tracking, on the same VRAM wall as the rest of video. And keep the decode last: face repair and HD upscale chain either way round, because both hand back a standard H3 latent, but MiniMax H3 Chunked Decode has to be the end of the line - that's where the RGB composite happens.
Inputs (18)
| Name | Type | Default | Description |
|---|---|---|---|
| sampled_av_latent | LATENT | — | |
| refine_context | STAR7_H3_REFINE_CONTEXT | — | |
| enable_refine | BOOLEAN | true | — |
| face_detector | COMBO | face_yolov8m.pt | Face detector in models/ultralytics/bbox. The pinned default downloads automatically when missing. |
| face_lora | COMBO | 继承一采 | Inherit the first-pass model unchanged, or apply the selected model-only LoRA for face repair only. |
| face_lora_strength | FLOAT | 1.00-100–100 | Model strength for the selected face-repair LoRA. |
| face_attention | COMBO | 继承一采 | 20 options: 继承一采, existing, comfy_kitchen_int8, sla_sm75_qk_int8_pv_fp16, sla_sm75_all_int8, sol_sm75_all_int8, +14 |
| face_count | INT | 11–4 | — |
| preset | COMBO | 自动平衡 | 5 options: 自动平衡, 真人保真, 远景小脸, 动漫角色, 自定义 |
| target_face | COMBO | 主人物 | 3 options: 主人物, 画面中央, 参考图匹配 |
| refine_steps | INT | 41–12 | — |
| custom_strength | FLOAT | 0.300.01–0.8 | Legacy H3 simple-scheduler denoise strength. Higher values can change identity and motion. Actual Sigma is shown in the report. |
| custom_canvas | COMBO | 自动 | 3 options: 自动, 512, 768 |
| custom_crop_context | FLOAT | 2.61.4–4 | — |
| custom_blend | FLOAT | 0.900–1 | — |
| custom_feather | INT | 200–64 | — |
| seed | INT | 00–18446744073709550000 | — |
| preserve_repair_detail | BOOLEAN | true | Composite in RGB after Star7 decode. When enabled, enlarge outputs below about 1 MP before pasting faces; when disabled, keep source resolution. Neither mode re-encodes the full video. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| refined_av_latent | LATENT | — |
| report | STRING | — |