Nodes/comfyui-minimax-h3-audio-T8/H3 Face Window · Bind Separate Stage (T8 EXP)
ComfyUI Node

H3 Face Window · Bind Separate Stage (T8 EXP)

Repairing one face in one stretch of frames, without re-rendering the film

By T8mars·Created 2 months ago·Updated about 7 hours ago· 1,158
H3 Face Window · Bind Separate Stage (T8 EXP)
  • face_plan
  • source_frames
  • parent_frames
  • window_plan
  • window_mapping
  • source_audio
  • window_audio
  • model
  • sampler
  • sigmas
  • av_latent
  • model
  • sampler
  • sigmas
  • stage_context
  • report_json
◄audio_policyrequire_locked►

The least glamorous and most useful face workflow: your 8-second clip is fine except the two seconds where someone turns their head and the profile goes soft. You do not want a full-clip refinement pass, because the other six seconds are correct and every extra pass is another chance to lose the likeness. You want to extract that window, fix it, and put it back.

That is what the window variant does. It is the region-and-time-scoped version of this pack's face pipeline, and this node is where the extracted window gets bound to a sampling stage.

What it binds

Five things have to agree, and the node checks all of them: the parity-style plan, the source frames, the parent frames, the signed window plan, and the window mapping. The window plan and mapping are the objects that say which absolute frames were extracted and how they map back to the parent - a signed contract rather than a remembered offset.

You also hand it the audio pair: source_audio for the whole clip and window_audio for the extracted stretch. Then the sampling environment: model, sampler, sigmas, av_latent.

audio_policy is the same choice as elsewhere in this pack. require_locked (default) insists on an exactly all-zero nested audio mask, i.e. the audio region is genuinely pinned; anything else errors and tells you so. preserve_existing takes the mask as it comes, which is the honest option when the audio is in scope. Getting this wrong does not corrupt anything silently - it refuses - but it is the field people stare at longest.

Outputs: the bound model, sampler, sigmas, a typed stage_context, and report_json.

What it deliberately does not do

It does not review, accept, or commit your repair. The pack's manual review and studio commit nodes stay outside this stage, and the description is explicit that there is no automatic acceptance or delivery. That is a design choice with teeth: the pipeline can produce a repaired window and stop, waiting for you to look at it and decide.

It also does not sample. Bind, sample, audit, review - four steps, and each one can fail loudly on its own. Annoying when you just want the face fixed; very good when the face comes back wrong and you need to know whether the plan, the sampler or your eyes were the problem.

Wiring

Window plan + mapping + parent frames + audio → Window Bind → parity sampling stage → Window Stage Audit → crop decode → manual/studio review → commit.

Installing it

cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8

Fully restart ComfyUI after cloning; registration happens at import and a refresh will not do it. Or install through ComfyUI Manager by searching "MiniMax H3 Audio T8". The pack's requirements.txt is empty on purpose - no pip step, no risk of this pack replacing ComfyUI's Torch or CUDA - and optional EXP features load their own requirements only when used. Recent Core with native H3 support plus H3 weights, Qwen encoder, video VAE and audio VAE are prerequisites.

Things that will bite you

Window work multiplies bookkeeping. The window mapping is signed - change the parent clip, or re-extract with a different range, and the existing mapping is invalid. Rebuild it rather than nudging numbers until the error goes away; the error is the point.

Second, keep the two audio inputs straight. Swapping source and window audio is the kind of mistake that produces a plausible-looking report and a wrong result, since both are audio tensors of different lengths and nothing intrinsic stops you.

Third, remember the window is a crop of a generated clip. The neighbouring frames are the model's own output, and neither this node nor the audit can tell you whether the repaired window matches its neighbours visually. That is what the manual review node is for, and it is worth actually using it - the pack's whole stance is that mechanical validation and human judgement are different things, and it keeps them in separate nodes on purpose.

CategoryT8/MiniMax H3/Modular Sampling/Experimental

Inputs (12)

NameTypeDefaultDescription
face_planH3_T8_FACE_REFINE_PARITY_PLAN—
source_framesIMAGE—
parent_framesIMAGE—
window_planH3_T8_FACE_REFINE_WINDOW_PLAN—
window_mappingH3_T8_FACE_REFINE_WINDOW_MAPPING—
source_audioAUDIO—
window_audioAUDIO—
modelMODEL—
samplerSAMPLER—
sigmasSIGMAS—
av_latentLATENT—
audio_policyCOMBOrequire_locked2 options: require_locked, preserve_existing

Outputs (5)

NameTypeDescription
modelMODEL—
samplerSAMPLER—
sigmasSIGMAS—
stage_contextT8_STAGE_CONTEXT—
report_jsonSTRING—