H16-3 · Explicit PASS2 Plan (T8 EXP)
Split your H3 second pass into windows before you burn 40 minutes on one 124-frame run
- first_pass_latent
- plan
- report_json
MiniMax H3 generates video and audio in one pass, so "upscaling" it isn't a post-process you bolt on afterwards - the 4+4 route in this pack runs four low-resolution steps, pushes the latent through a learned 3D upscaler, then runs four more steps at the higher size with a fresh noise field. That second pass is the expensive one, and on a 124-frame clip it wants the whole timeline resident at once.
This node is the escape hatch. It takes the already-completed HIGH latent and writes out an explicit plan for running that second pass one temporal window at a time, so a long clip becomes seven short jobs instead of one job that eats your VRAM and dies at 70%.
What it actually does
Plan node. It does not sample, load models, or touch the GPU. It reads the geometry off your first_pass_latent (the video latent from the completed first pass), multiplies the latent's width/height by the H3 VAE downsample factor to get the real pixel target, and builds the same full-frame chunked two-pass plan the old all-in-one H16 node used internally - the "v4" plan from the learned 3D upscale path.
Spatial strategy is pinned to full_frame_safe, meaning every window gets the full canvas. There is no secret tiling. Temporal windows are the only axis you control, and the plan carries a SHA that every downstream window node re-checks: the whole point is that window 3 must be provably sampled against the exact same plan window 0 used.
The inputs you'll actually touch
Two of them matter on day one.
temporal_strategy defaults to guarded_overlap_exp, which chunks the timeline and keeps a guarded takeover in the overlap region so the seam doesn't pop. Set it to full_clip_safe if you want the old behaviour back and one giant window - which is the same thing as not using this route at all.
temporal_chunk_frames (default 34) and temporal_overlap_frames (default 17) both step in 17s, because H3's temporal grid is 17n. Don't invent a 30-frame window; you'll get a plan whose window count can't line up with the number of PASS2 nodes you wired. The classic example config is 124 frames = seven 34-frame windows with 17 frames of overlap.
anchor_strength (default 0.999) is how hard each window is anchored to its context; leave it alone unless you're deliberately chasing a looser motion feel. model_name defaults to minimax_h3_latent_upscaler_3d_fp16.safetensors and must be a real file in ComfyUI/models/latent_upscale_models/ - the plan is written against that specific upscaler's geometry.
Outputs: plan (typed, goes to your chunked source/prepare nodes and every window node) and report_json, which is worth reading once. It tells you how many windows you just created, which is how many PASS2 nodes you owe the graph.
Installing it
Same for every node in this pack. ComfyUI Manager → search MiniMax H3 Audio T8, or:
cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8
Restart ComfyUI fully (not just refresh). The pack's requirements.txt deliberately installs nothing - torch, torchaudio, numpy, Pillow and safetensors come from ComfyUI, and the author keeps them out so the pack can't clobber your CUDA stack. You do need a recent ComfyUI core with native H3 support, plus the H3 base model, Qwen3-VL text encoder and the video/audio VAEs. One licence note since it's real: the H3 Community License excludes the US, EU, UK and South Korea from its territory, so the local weights are licensed out for a lot of readers.
Where people get burned
The plan is a contract, not a suggestion. Change temporal_chunk_frames or anchor_strength after you've already sampled window 1 and the next window node will refuse with a stale-plan error rather than quietly blending two different plans. And the window count is fixed at plan time: if the report says seven windows and you wired four, you are simply not going to get the last half of your clip.
Also: this node is useless on its own. It needs the chunked source/prepare nodes feeding source_segment/segment_spec/pass2_context and the learned lift output feeding lifted_segment on each window. If your second pass still runs fine as one shot, keep it that way - windowing costs you a mountain of nodes and only pays for itself at long durations.
Inputs (6)
| Name | Type | Default | Description |
|---|---|---|---|
| first_pass_latent | LATENT | — | |
| temporal_strategy | COMBO | guarded_overlap_exp | 2 options: guarded_overlap_exp, full_clip_safe |
| temporal_chunk_frames | INT | 3417–3600 | — |
| temporal_overlap_frames | INT | 170–1700 | — |
| anchor_strength | FLOAT | 0.9990–1 | — |
| model_name | STRING | minimax_h3_latent_upscaler_3d_fp16.safetensors | — |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| plan | T8_H3_CHUNKED_TWO_PASS_PLAN | — |
| report_json | STRING | — |