Nodes/comfyui-minimax-h3-audio-T8/H3 Continuation · Geometry + Stage Plan (T8 EXP)
ComfyUI Node

H3 Continuation · Geometry + Stage Plan (T8 EXP)

The node that decides which 4 of your 8 steps run cheap

By T8mars·Created 2 months ago·Updated about 7 hours ago· 1,158
H3 Continuation · Geometry + Stage Plan (T8 EXP)
  • contexts
  • full_sigmas
  • plan
  • low_sigmas
  • high_sigmas
  • length
  • report_json
◄length124►
◄low_evaluations4►
◄low_scale0.50►

MiniMax H3 is the 33B open-weights video model that generates picture and stereo audio in one pass, and ComfyUI has had native support since day one. The catch for anyone pushing past a single clip is that H3 was not trained to run forever: long output means stitching segments, and stitching segments means deciding how much of each new segment you pay full price for. This node is that decision, written down as geometry.

What this node is for

The pack ships two long-video philosophies. The old one is a single all-in-one node that loops segments internally, and it works. The modular route - the EXPT8 continuation nodes - does the same job with the stages pulled apart: pick an accepted parent, prepare its context, plan the geometry, sample low resolution, hand the result to a learned 3D latent upscaler, then sample the high-resolution tail. Nothing is hidden, which is the whole point. If the seam looks wrong you can point at which stage produced it.

Plan is the second step of that chain and the only node that touches timing. It does not read a prompt, does not see a model, and cannot sample anything.

What it computes

Two things: an accepted-parent geometry and one native Euler schedule, split in half.

You hand it the full sigma table your native H3 setup already produces via full_sigmas. It splits that table at low_evaluations and hands you back both halves - low_sigmas is the first low_evaluations + 1 values, high_sigmas is everything from low_evaluations onward. An eight-step table with low_evaluations at 4 gives you the familiar 4+4: four steps drawn at the low canvas, the learned 3D upscaler expands the latent, four more steps at final resolution. It is not two separate four-step generations, and the HIGH half does not restart its clock.

The geometry half is about the accepted clip you are continuing. Frame lengths live on the same 17n+5 grid H3's long-video path uses - 5, 22, 39, … 124, 141 - which is why length steps by 17 instead of 1. Put 100 in there and the downstream stages will not have a legal window to build.

The inputs that matter

  • contexts - the typed context object from the "Select Accepted Parent" and "Prepare Accepted Contexts" pair. It carries the parent's exact frame count, canvas and low-resolution canvas, so the plan can validate everything against it.
  • full_sigmas - from your native H3 setup. One table, shared by both halves.
  • length - render window for this segment, default 124, must sit on the 17n+5 grid.
  • low_evaluations - how many steps run at reduced scale. Default 4. Must leave at least one step for HIGH.
  • low_scale - spatial scale of the low canvas, default 0.5, aligned to 32 pixels. 0.5 on 896×448 lands on 448×224.

Out comes plan, which is the object every later stage verifies against; low_sigmas and high_sigmas for the two samplers; length passed through as an integer for wiring convenience; and report_json, which is the receipt you paste into a bug report when something refuses to run.

Wiring, and the one rule that matters

Plan's plan output feeds the LOW sampler stage, the HIGH handoff stage and the Relay apply node. That shared object is the contract: the HIGH handoff will reject a boundary that does not come from this plan, and the Relay nodes will reject a window whose frame count does not match the encoded phase.

The rule: connect the common progressive lift input and an external learned 3D latent upscaler. The author is explicit that this is not FastH3, HyperFlow or VDN math - those are separate acceleration routes with their own schedules. Mixing a HyperFlow split or a FastH3 distilled table into this plan is not "combining optimisations", it is feeding a schedule to code that did not expect it.

Installing it

Everything in this pack installs the same way:

cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8

Then fully quit ComfyUI and restart - not a browser refresh, an actual process restart, because node registration happens at import. ComfyUI Manager can also find it by searching the pack title "MiniMax H3 Audio T8".

requirements.txt in this pack deliberately installs nothing; the base nodes ride on ComfyUI's own torch, numpy, Pillow and safetensors, and optional Advanced/EXP features check their own dependencies only when you actually use them. That is good news for dependency hell, which is the usual way a video node pack breaks your environment. You still need a recent ComfyUI with native H3 support, and you still need the H3 weights, Qwen encoder and both VAEs in their usual folders.

Things that will bite you

The node will refuse to run before it produces anything wrong: mismatched contexts, a length off the 17n+5 grid, or a plan object rebuilt from a different parent all raise. That is the design - the author's whole release posture is "reject rather than silently re-infer".

reserve_vram_mib on the sampling stages is a boundary headroom check, not a ceiling on your card, and the docs say plainly it does not guarantee you avoid an OOM. If 32-second windows die on a 16GB card, lowering the reserve does not fix the actual working set.

CategoryT8/MiniMax H3/Modular Sampling/Continuation Experimental

Inputs (5)

NameTypeDefaultDescription
contextsT8_CONTINUATION_STAGE_CONTEXTS—
full_sigmasSIGMAS—
lengthINT124—
low_evaluationsINT41–999—
low_scaleFLOAT0.500.25–0.99—

Outputs (5)

NameTypeDescription
planT8_PROGRESSIVE_STAGE_PLAN—
low_sigmasSIGMAS—
high_sigmasSIGMAS—
lengthINT—
report_jsonSTRING—