FastH3 V2 · Accepted Joint-Audio Reconcile (T8 EXP)
Picking which half's audio survives a 4+4 continuation
- contexts
- learned_latent
- highres_template
- positive
- reconciled_av
- positive
- report_json
H3 generates picture and stereo sound in the same pass, which is lovely until you split the pass in two. The old FastH3 V2 continuation recipe is 4+4: four distilled steps at low resolution, a learned latent upscale, four more steps at final resolution. Each half produces an audio stream. Only one of them can be in your file, and choosing badly is how a smooth-looking clip ends up with a two-second audio glitch at the seam.
This node makes that choice explicit, as a step you own rather than a policy hidden inside a bigger node.
What it does
It reproduces the legacy V2 continuation audio selection - the one the original all-in-one 4+4 node performed - from two inputs: the learned low-resolution x0 and the fresh high-resolution template. Out comes reconciled_av, an AV latent with the video from the high-resolution path and the audio selected per that legacy policy. It also passes positive back out, so you can keep chaining conditioning.
Two practical points from the author's description deserve highlighting. There is no hidden sampler in here, and there is no audio-freeze shortcut: the node is choosing between two genuinely generated audio streams, not duct-taping a frozen copy of the first pass over the top of the second.
Inputs and outputs
contexts- the authenticated accepted window, same object the other V2 ports take. It is the identity anchor; the reconcile result is tied to it.learned_latent- the low-resolution denoisedx0after the external learned latent upscale step. Note the ordering that implies: your upscaler runs before this node, not after it.highres_template- the fresh high-resolution AV latent layout at the target canvas, i.e. the packed AV your high-resolution pass was built against - a template, not a sampled result.positive- conditioning passed through and returned.
Outputs are reconciled_av and positive, plus report_json.
Ordering, because this is where people go wrong
The HIGH prefix node goes after this one. The sequence the pack documents is:
- LOW pass, 4 distilled steps at low resolution.
- External learned 3D latent upscale.
- Reconcile - video from the high-res path, audio selected per the legacy policy.
- HIGH prefix lock - apply the continuation prefix and optional mask ramp to the reconciled AV.
- HIGH pass, 4 more steps at final canvas.
Doing the prefix first and reconciling afterwards throws away the thing reconcile is supposed to preserve. The pack's own implementation asserts that the prefix step keeps the same reconciled audio tensor, which is a machine-checkable statement that the order matters.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8
Then fully restart ComfyUI. Registration happens at import; a browser refresh will not make the node appear. Manager users: search "MiniMax H3 Audio T8". The pack installs no pip dependencies by design - no chance of it swapping ComfyUI's Torch or CUDA stack, and the V2 student does not need FastVideo's distributed runtime. You do need the pinned V2 checkpoint in models/diffusion_models and the usual Qwen encoder plus video and audio VAEs.
Things that will bite you
If your seam audio is wrong, the first suspect is the order, not the reconcile node. The second is which latent you fed in: learned_latent must be the low-resolution x0 after the learned upscale, and highres_template must be the high-resolution layout - swapping or bypassing either produces a reconciled AV that looks plausible and is built on the wrong stream.
Third, the general warning the pack repeats everywhere: reproducing a legacy policy is not a claim that the result is good. The acceptance evidence in this pack is bound to specific samples, models and configurations, and it explicitly does not generalise to arbitrary reference images, LoRAs or long windows. It also does not cover combining the trained VSA attention path with Relay timing on V2 - the pack documents that combination as unsupported rather than silently degraded. Run your own segment, watch the seam, listen to the whole audio, then queue the long job.
Inputs (4)
| Name | Type | Default | Description |
|---|---|---|---|
| contexts | T8_CONTINUATION_STAGE_CONTEXTS | — | |
| learned_latent | LATENT | — | |
| highres_template | LATENT | — | |
| positive | CONDITIONING | — |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| reconciled_av | LATENT | — |
| positive | CONDITIONING | — |
| report_json | STRING | — |