Nodes/MiniMax H3 Activation Chunk - Star7/MiniMax H3 One-click Face Repair - Star7
ComfyUI Node

MiniMax H3 One-click Face Repair - Star7

Your H3 clip is fine except the face — this node fixes just the face

By star7code·Created 2 months ago·Updated 5 days ago· 86
MiniMax H3 One-click Face Repair - Star7
  • sampled_av_latent
  • refine_context
  • refined_av_latent
  • report
◄enable_refinetrue►
◄face_detectorface_yolov8m.pt►
◄face_lora继承一采►
◄face_lora_strength1.00►
◄face_attention继承一采►
◄face_count1►
◄preset自动平衡►
◄target_face主人物►
◄refine_steps4►
◄custom_strength0.30►
◄custom_canvas自动►
◄custom_crop_context2.6►
◄custom_blend0.90►
◄custom_feather20►
◄seed0►
◄preserve_repair_detailtrue►

You spent twenty minutes on a clip and the shot is great - motion, framing, the audio syncs - and the face looks like it was painted with a wet brush. In image land you'd reach for an ADetailer-style pass: detect, crop, re-diffuse the crop, paste it back. Video gives you no such free lunch - what you have after sampling is a packed H3 latent holding video and audio as one nested tensor.

MiniMaxH3FaceRefineLatentStar7 (MiniMax H3 One-click Face Repair - Star7) is that detailer pass, built for H3's latent layout.

Why it needs a second cable

A second sampling pass needs your model chain, your positive conditioning and the video VAE - without you re-wiring the loader, the LoRAs and the chunk node a second time. That's what refine_context is: a STAR7_H3_REFINE_CONTEXT bundle emitted by MiniMax H3 All-in-one Conditioning - Star7 (same pack). Feed it anything else and it stops with Star7 face-repair context is incomplete. So the swap is: all-in-one conditioning instead of your usual one, plus a single wire.

All-in-one Conditioning "refine_context" --> Face Repair "refine_context"
Sampler "sampled_av_latent" --------------> Face Repair "sampled_av_latent"
Face Repair "refined_av_latent" ----------> Optional HD Upscale --> Chunked Decode

How the pass actually works

The node decodes the video member of the packed latent to frames, runs YOLO face detection with per-frame tracking (with automatic scene-cut detection, so tracks don't smear across cuts), and crops each track onto a square canvas - auto caps it at 768, or you can pin 512/768. It then builds a fresh H3 video latent at canvas size, keeps the original audio latent, applies a per-frame denoise mask, and samples just that crop: simple schedule, res_multistep, refine_steps steps.

The part worth understanding, because it changes how you wire the end of the graph: the refined crops are not decoded and re-encoded into the full video. The node packs each repaired crop into a compact overlay carrying the tracking transform and your blend/feather settings, attaches it to the latent, and hands back your original sampled_av_latent untouched - base video latent and audio preserved exactly. The RGB stitch happens once, later, in MiniMax H3 Chunked Decode.

A warning as much as a design note: a stock VAEDecode at the end of your graph ignores the overlay and hands you the original, un-repaired footage.

Inputs you'll actually touch

  • sampled_av_latent - the sampler's output, and it has to be the real thing. A plain image latent gets you requires a sampled MiniMax H3 packed audio-video LATENT.
  • refine_context - from the all-in-one conditioning node, as above.
  • preset - the four named presets are tuned, not decorative; they set how hard the crop gets redrawn. 自动平衡 starts at 0.30 denoise, 真人保真 is gentler (0.25, tighter crop, 512 canvas), 远景小脸 is the wide-shot preset (0.48 denoise, 768 canvas), 动漫角色 sits at 0.32. The custom_* fields only apply when preset is 自定义 - the most common "why did nothing change" trap here.
  • face_count (1–4) and target_face - 主人物 tracks the largest face, 画面中央 the centre-most, and 参考图匹配 needs insightface (otherwise it warns and falls back to main-subject tracking).
  • refine_steps (1–12), seed, preserve_repair_detail - the last one defaults on and lifts video below about 1 MP to a 0.98–1.01 MP canvas before pasting, so the repaired crop isn't squashed on the way back in. It never downscales.

face_lora, face_lora_strength and face_attention all default to inheriting the first pass. Point face_attention at another backend and it applies to the repair pass only, then gets restored - handy when you want the faces running under a different backend than the bulk video.

Outputs and installing

Two outputs: refined_av_latent (LATENT) and report (STRING). Read the report: final resolution, the sigma the pass started from, the detector, the LoRA and attention in play - the fastest way to tell "it ran but the effect is subtle" apart from "it didn't run".

Install via Manager (search MiniMax H3 Activation Chunk - Star7), comfy node install minimax-h3-chunk-star7, or:

cd ComfyUI/custom_nodes
git clone https://github.com/star7code/minimax-h3-chunk-star7.git

The pack pulls in ultralytics (the detector runtime) and scenedetect. The face_detector dropdown defaults to face_yolov8m.pt; on first use the node downloads it from Bingsu/adetailer - HF mirror first, then HF - verifies the SHA-256 and drops it in ComfyUI/models/ultralytics/bbox. The same dropdown lists any other face detector you already have in that folder.

Where people get burned

It no-ops instead of erroring when it can't find a face. no face detected and no stable face track both log a warning and pass your latent straight through, so the report string is where you find out. For a wide shot with a genuinely tiny face, try 远景小脸 before deciding the node is broken.

Budget for it. It decodes the source video up front and runs one extra sampling pass per face - face_count=4 is four passes plus detection and tracking, on the same VRAM wall as the rest of video. And keep the decode last: face repair and HD upscale chain either way round, because both hand back a standard H3 latent, but MiniMax H3 Chunked Decode has to be the end of the line - that's where the RGB composite happens.

CategoryStar7/MiniMax H3

Inputs (18)

NameTypeDefaultDescription
sampled_av_latentLATENT—
refine_contextSTAR7_H3_REFINE_CONTEXT—
enable_refineBOOLEANtrue—
face_detectorCOMBOface_yolov8m.ptFace detector in models/ultralytics/bbox. The pinned default downloads automatically when missing.
face_loraCOMBO继承一采Inherit the first-pass model unchanged, or apply the selected model-only LoRA for face repair only.
face_lora_strengthFLOAT1.00-100–100Model strength for the selected face-repair LoRA.
face_attentionCOMBO继承一采20 options: 继承一采, existing, comfy_kitchen_int8, sla_sm75_qk_int8_pv_fp16, sla_sm75_all_int8, sol_sm75_all_int8, +14
face_countINT11–4—
presetCOMBO自动平衡5 options: 自动平衡, 真人保真, 远景小脸, 动漫角色, 自定义
target_faceCOMBO主人物3 options: 主人物, 画面中央, 参考图匹配
refine_stepsINT41–12—
custom_strengthFLOAT0.300.01–0.8Legacy H3 simple-scheduler denoise strength. Higher values can change identity and motion. Actual Sigma is shown in the report.
custom_canvasCOMBO自动3 options: 自动, 512, 768
custom_crop_contextFLOAT2.61.4–4—
custom_blendFLOAT0.900–1—
custom_featherINT200–64—
seedINT00–18446744073709550000—
preserve_repair_detailBOOLEANtrueComposite in RGB after Star7 decode. When enabled, enlarge outputs below about 1 MP before pasting faces; when disabled, keep source resolution. Neither mode re-encodes the full video.

Outputs (2)

NameTypeDescription
refined_av_latentLATENT—
reportSTRING—