H3 Chunked v5 · One Joint AV PASS2 Window (T8 EXP)
The node that actually renders your chunked H3 windows
- model
- positive
- partial4_denoised_output
- lifted_full_av
- prepared
- plan
- noise
- sampler
- sigmas
- previous_result
- negative
- cumulative_av_latent
- window_result
- report_json
This is the business end of the v5 chunked family. H3 Chunked v5 · One Joint AV PASS2 Window runs exactly one "remaining 4" refine window over the globally-lifted AV latent, with MODEL, CONDITIONING, NOISE, SAMPLER and SIGMAS all wired in from outside. Then you instantiate it again for the next window and chain window_result.
If you've used Wan 2.2's high-noise/low-noise two-stage setup, the shape of the recipe will feel familiar: a cheap first half, an upscale, a refine half. What's different is that H3's audio lives in the same latent, so "joint AV" here means the refine pass is shaping picture and sound together, and the audio in the final clip is the refined joint audio rather than a passthrough of the first pass.
What you wire in
Required: model, positive, partial4_denoised_output, lifted_full_av, prepared, plan, noise, sampler, sigmas, window_index (0-based, default 0) and cfg (default 1, step 0.1 - just leave it at 1). Optional: previous_result, and negative (near-irrelevant at CFG 1, but harmless).
Out: cumulative_av_latent - the whole clip as it stands after this window, which is what you decode at the end. window_result - the typed T8_CHUNKED_V5_WINDOW_RESULT the next window wants as previous_result, and which the EAV/Relay apply and the freeze/load nodes consume. And report_json.
How the stitching stays honest
Later windows read the previous window's completed overlap as read-only context: the mask over that region is zeroed, and after sampling the node checks that the overlap's values are unchanged, element by element. Native Euler's last float32 add can wobble by a few ULPs, so it allows a bound of roughly 8 * float32_eps * (1 + |source|) and then restores the masked values exactly. Anything bigger and it raises. There is no averaging, no crossfade in RGB, no latent endpoint slide and no global "close enough" threshold to hide real drift.
Practical consequence: the window count and overlap are properties of the plan, and you should not try to hand-tune a window's boundaries to paper over a content shift. Bigger overlap means more compute and more independent prediction trajectory, not less drift, and the pack says so in as many words.
Sizing, from the pack's documented run: 192 frames at 24 fps (8 s), 136-frame window, 34-frame overlap → windows [0,136) and [102,192). That's the shared four-step low pass plus four steps per window: 12 model forwards, not 8. Single-window clips are the only ones where "4+4 = 8" is literally true.
Install
Manager → search MiniMax H3 Audio T8, or:
cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8
Fully restart ComfyUI, then refresh the tab. You need a recent Core with native H3 support, plus the H3 main model in models/diffusion_models, Qwen in models/text_encoders, and the video and audio VAEs in models/vae. No extra pip packages - that empty requirements.txt is intentional. The author's example workflows live in the repo's examples/workflows tree; for v5 that's the 54-chunked-v5-split set, which includes per-window save and cold-resume variants.
Where it goes wrong
Skipping the prepare node. prepared is not optional and not interchangeable between runs. All windows of one clip share one preparation; wiring a second prepare in for window 2 is the classic mistake.
Forgetting previous_result. Window 1 will run happily with no chaining and then refuse to merge, or refuse outright. The typed result is the handoff, not a suggestion.
Reading the wrong output. cumulative_av_latent is the accumulating whole; window_result is bookkeeping. Decode and save the cumulative one - and note the pack explicitly says cumulative intermediates are not portable cache receipts. For a resumable artifact use Freeze ONE Joint AV Window.
The green-node trap. This is an EXP node: it can be wired perfectly and still be the wrong tool for your material. The README's own boundary applies - EXP contracts are qualified on specific samples, seaming still needs human review, and there's a known slight seam colour shift on dual-pass stitching.
Stale duplicate pack. Missing sockets or absent nodes after an update usually mean an old copy of this pack in custom_nodes shadowing the real one. Rename the stale directory with a .disabled suffix (a leading underscore won't disable it), restart, and check python_module in /object_info.
Inputs (13)
| Name | Type | Default | Description |
|---|---|---|---|
| model | MODEL | — | |
| positive | CONDITIONING | — | |
| partial4_denoised_output | LATENT | — | |
| lifted_full_av | LATENT | — | |
| prepared | T8_CHUNKED_V5_PREPARED | — | |
| plan | T8_H3_CHUNKED_TWO_PASS_PLAN | — | |
| noise | NOISE | — | |
| sampler | SAMPLER | — | |
| sigmas | SIGMAS | — | |
| window_index | INT | 00–9999 | — |
| cfg | FLOAT | 1.00–100 | — |
| previous_resultopt | T8_CHUNKED_V5_WINDOW_RESULT | — | |
| negativeopt | CONDITIONING | — |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| cumulative_av_latent | LATENT | — |
| window_result | T8_CHUNKED_V5_WINDOW_RESULT | — |
| report_json | STRING | — |