H3 Face · Bind Separate Refinement Stage (T8 EXP)
Give the face repair its own sampler instead of hijacking the whole video pass
- face_plan
- source_frames
- model
- sampler
- sigmas
- av_latent
- model
- sampler
- sigmas
- stage_context
- report_json
The first thing to understand about face repair in this pack is that it happens after the video is generated, as a separate low-denoise pass over a face crop. You generate the clip with H3, notice the face went soft or slightly off-model halfway through, and then you fix the face without regenerating the film and without touching the audio.
That means the repair has its own model, its own sampler, its own sigma table and its own conditioning. This node is the adapter that assembles those four things and checks that they actually belong to the face plan you are repairing.
What it does
Face repairs in this pack come from a plan - a signed object built by the pack's face-sourcing nodes that knows the source clip, the frame count, the canvas, the crop geometry and which frames are in scope. The bind node takes that plan, the source frames, your model/sampler/sigmas and the AV latent, and checks them against each other:
- the source frames' frame count and dimensions must match what the plan recorded;
- the AV latent's time and canvas must match the plan's aligned geometry;
- the audio policy must be satisfiable.
Only when all of that agrees does it return the stage, plus a typed stage_context for the sampling stage.
No sampling, no VAE encoding, no stitching. It is the "wire it up correctly" half of a two-node pattern; the audit node is the other half.
The inputs that matter
face_plan- the standard face plan. There are variants in this pack (a parity plan for per-frame denoise, a window plan for a cropped repair window, a multi-person plan), and each has its own bind node. Handing a parity plan to this node is a mistake that errors rather than a configuration that half-works.source_frames- the actual frames being repaired. Not the parent clip's file, the IMAGE tensor. Mismatch here is the single most likely error you will see, and it means your plan was built for a different clip, a different frame count, or a different canvas.model,sampler,sigmas- the repair stack. The intended shape is a low-denoise pass: a small sigma range, so the model is cleaning and re-detailng rather than inventing a new face.av_latent- the packed audio/video latent being repaired.audio_policy- two options, defaultrequire_locked. Inrequire_lockedthe nested audio mask must be present and exactly all-zero, i.e. the audio region is genuinely pinned and nothing will be resampled there. If your AV latent has a non-zero audio mask, this errors, and the fix is either to lock the audio properly or to selectpreserve_existingif you know you are regenerating sound. Choosingrequire_lockeddeliberately is how you avoid an audio pass quietly changing the soundtrack under a "video only" repair.
Outputs: model, sampler, sigmas (the bound stack, which is what you actually hand to the sampling stage), stage_context (the typed context the stage sampler and the audit node need), and report_json.
Wiring
Face plan + source frames → Bind → stage sampler → audit → existing stitch node.
The bind node contributes no outputs to the final video by itself. If you have wired it and nothing else, you have wired half a pipeline.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8
Restart ComfyUI fully - these classes register at import, so a page refresh will not show them. ComfyUI Manager can install it too; search "MiniMax H3 Audio T8". The pack ships an empty requirements.txt on purpose, so installing it cannot swap ComfyUI's Torch or CUDA stack - a real mercy in a video-node ecosystem where one pack's pin on transformers can break another's. You still need a recent Core with native H3 support, H3 weights, Qwen encoder, and both VAEs.
Things that will bite you
Read the error text before touching the workflow. "Face source dimensions differ from the bound plan" means the plan and the frames disagree - rebuild the plan from the current clip rather than resizing something to fit. "Face require_locked needs an exactly zero nested audio mask" means your audio is not actually locked and you asked for it to be.
And resist the temptation to push repair strength up until the face looks perfect. Everything in this pack that gains strength also gains freedom to drift, and a face that is sharper but no longer the person is a worse outcome than a soft one. Start small, compare against the unfixed frames, and keep the source audio out of the repair entirely - the pack's own audit node notes that sampled audio is not the delivered source track.
Inputs (7)
| Name | Type | Default | Description |
|---|---|---|---|
| face_plan | H3_T8_FACE_REFINE_PLAN | — | |
| source_frames | IMAGE | — | |
| model | MODEL | — | |
| sampler | SAMPLER | — | |
| sigmas | SIGMAS | — | |
| av_latent | LATENT | — | |
| audio_policy | COMBO | require_locked | 2 options: require_locked, preserve_existing |
Outputs (5)
| Name | Type | Description |
|---|---|---|
| model | MODEL | — |
| sampler | SAMPLER | — |
| sigmas | SIGMAS | — |
| stage_context | T8_STAGE_CONTEXT | — |
| report_json | STRING | — |