Nodes/comfyui-minimax-h3-audio-T8/FastH3 V2 · Accepted HIGH Prefix After Reconcile (T8 EXP)
ComfyUI Node

FastH3 V2 · Accepted HIGH Prefix After Reconcile (T8 EXP)

Lock the old segment in before the distilled tail runs

By T8mars·Created 2 months ago·Updated about 7 hours ago· 1,158
FastH3 V2 · Accepted HIGH Prefix After Reconcile (T8 EXP)
  • contexts
  • reconciled_av
  • high_av
  • report_json
◄modehigh_native_mask_ramp_exp►

Continuing a video is easy to describe and hard to get right: the new segment has to start from the old one's last frames, but those frames must not be re-generated. If the sampler is allowed to touch them, it invents a slightly different version of a shot you already approved, and the seam inherits the difference. Every serious long-video workflow has some mechanism for pinning that prefix. This is the FastH3 V2 one, exposed as its own node.

What it does

It applies the old V2 continuation HIGH video prefix to a reconciled AV latent, with an optional three-step mask ramp, and returns the result as high_av - which is what you feed the high-resolution distilled pass.

The prefix portion is the accepted completed context: the last frames of the parent, carried into the new window's video channels. The ramp is the transition - a short, stepped fade of the mask across the first few frames of the new material rather than a hard cut in the latent, so the distilled sampler has somewhere to blend instead of an edge to resolve.

One thing it deliberately does not do: sample. The description says so plainly, and it matters for debugging, because it means a bad result here is a wiring or configuration problem, not a sampler behaving badly.

Inputs and outputs

  • contexts - the authenticated accepted window. The node checks the accepted context length against the native context steps and refuses if it is not a legal value.
  • reconciled_av - the AV latent from the joint-audio reconcile step. Ordering is fixed: reconcile first, prefix after. The implementation asserts that the prefix step preserves the same reconciled audio tensor, so doing it the other way round is not "an alternative workflow", it is discarding the audio decision you just made.
  • mode - high_native_mask_ramp_exp (default) applies the three-step mask ramp, high_native_mask_exp applies the native mask without the ramp.

Outputs are high_av for the high-resolution pass, and report_json with the segment index, context length and the legacy report fields.

Wiring, in order

LOW distilled pass → external learned 3D latent upscale → Reconcile → HIGH Prefix → HIGH distilled pass → decode.

Both this node and reconcile hang off the same contexts object, which is what keeps them talking about the same accepted parent. If you change the parent, both invalidate.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8

Full restart of ComfyUI afterwards - these classes register at import, so refreshing the page changes nothing. ComfyUI Manager can install it too; search "MiniMax H3 Audio T8". The pack's requirements.txt is deliberately empty, which is the single best thing about installing it: no pip run, no risk of it pinning over ComfyUI's Torch or CUDA. The FastH3 V2 student also does not require FastVideo's distributed runtime - just the checkpoint in models/diffusion_models, alongside the usual Qwen encoder and both VAEs.

Things that will bite you

The classic one: skipping the learned upscale and going straight from a low-resolution latent to the prefix. The prefix has to lock in video at the target canvas, so what feeds it has to already BE at that canvas - via the learned 3D latent upscaler this pack documents, not via a resized latent or a pixel upscaler. Wrong source, hard error.

Second, the context length. The accepted window carries a context frame count of 5, 22 or 39 - the native steps. A context that is not on that grid is rejected, which is the correct behaviour but confusing if you built the window with an arbitrary frame count upstream.

Third, check the mode before blaming the seam. The ramp is what makes a hard boundary look like a cut; with the non-ramp mode you are asking for the native mask and should expect a crisper transition. Neither mode changes the fact that this is the old V2 continuation behaviour being reproduced, not a new algorithm - and the pack is careful to say that reproducing it is not a quality claim about your material. Watch the boundary in motion before committing to a multi-segment chain.

CategoryT8/MiniMax H3/Modular Sampling/Continuation Experimental

Inputs (3)

NameTypeDefaultDescription
contextsT8_CONTINUATION_STAGE_CONTEXTS—
reconciled_avLATENT—
modeCOMBOhigh_native_mask_ramp_exp2 options: high_native_mask_exp, high_native_mask_ramp_exp

Outputs (2)

NameTypeDescription
high_avLATENT—
report_jsonSTRING—