Nodes/comfyui-minimax-h3-audio-T8/MiniMax H3 HyperFlow 双采长片 / Long Video (EXP/T8)
ComfyUI Node

MiniMax H3 HyperFlow 双采长片 / Long Video (EXP/T8)

Where the seam lands, and how resume works

By T8mars·Created 2 months ago·Updated about 7 hours ago· 1,158
MiniMax H3 HyperFlow 双采长片 / Long Video (EXP/T8)
  • model_pass1
  • model_pass2
  • clip
  • video_vae
  • audio_vae
  • video
  • video_path
  • manifest_path
  • completed_segments
  • status
  • report_json
◄hyperflow_file▾►
◄low_width448►
◄low_height224►
◄upscaler_model▾►
◄color_matchtrue►
◄chain_idh3_hyperflow_long_video_exp►
◄total_duration_seconds8.00►
◄width896►
◄height448►
◄render_window_frames124►
◄context_frames22►
◄global_prompt►
◄segment_prompts_json►
◄minimum_free_vram_mib512►
◄base_seed123456789►
◄seed_policyincrement►
◄resume_existingtrue►
◄filename_prefixH3_HyperFlow_Long_Video_EXP►
◄audio_seam_policycosine_bridge►
◄bridge_ms5.0►
◄bit_depth8►
◄crf18►

H3 clips are short by nature, so anything past a few seconds means continuation: generate a window, then drive the next window from what you already accepted. MiniMaxH3HyperFlowLongVideoEXPT8 is the whole 8-second, two-segment pipeline as one node - and unlike most "long video" nodes, it exposes the awkward parts as settings instead of hiding them.

One framing you need up front: 8 seconds here means two segments of a single 8-second film (124 + 68 frames, 192 total, ~5.17 s seam), not two 8-second halves.

What it does per segment

Each segment runs the two-stage HyperFlow upscale: LOW sampling over absolute intervals 0:4, learned 3D latent upscale, then HIGH intervals 4:8 with fresh noise. Both stages get their own patched model, which is why the node has two MODEL inputs: model_pass1 and model_pass2. They must come from one full H3 base, and hyperflow_file selects the original adapter separately - never through a generic LoRA loader.

The continuation rule is the interesting part. The next segment's LOW video condition is rebuilt from the last 39 frames of the previously accepted MP4 - resized and re-encoded through the VAE - while HIGH keeps the original high-resolution context and the audio continues from the completed HIGH stage. So audio is never frozen from a half-finished LOW; the node would rather keep the sound evolving than ship an artifact.

Then the seam dressings: color_match (on by default) and an audio seam policy defaulting to cosine_bridge with a 5 ms bridge_ms. The docs are careful about what these do - they don't repair a structural jump cut, and they aren't a human review.

Inputs and outputs worth knowing

Beyond the models: low_width / low_height (448×224), upscaler_model, color_match, clip, video_vae, audio_vae, and the geometry set - width/height (896×448), render_window_frames (124) and context_frames (22), which this node constrains hard because the recipe is shape-locked.

Prompting is either a global_prompt or segment_prompts_json for per-segment overrides. If Prompt Relay owns the timeline, both should be empty or exactly match the plan's global prompt; the tooltip spells that out because a mismatch is a silent-ish quality bug.

The reliability knobs are where this node earns its keep: chain_id namespaces the stage cache, base_seed plus seed_policy, resume_existing (on), minimum_free_vram_mib (an explicit start floor, rechecked before each segment - the tooltip warns it's not a peak guarantee), and filename_prefix.

Outputs: video, video_path, manifest_path, completed_segments, status, report_json. The manifest and the status/segment count are what you read to know whether a long run actually finished, instead of guessing from the output folder.

Resume is real, and it's where the value is

The pack's isolated tests interrupted a run mid-second-segment and resumed it with one HIGH 4-evaluation stage - first segment and the LOW/HIGH-input receipts reused, final MP4 byte-identical to the uninterrupted run. Discarding a stage copy and resuming from it returned the same file; a deliberately corrupted stage tensor was rejected before sampling.

That's why chain_id is the one field you must not reuse. Change the recipe - models, resolution, prompts, encoding - and you need a new chain id. The node validates frozen contracts (both bare model identities, HyperFlow file hash, the dual-clock grid and absolute intervals, upscaler hash, geometry, LOW size, seam and encode settings, source hash) and rejects bad or changed receipts rather than quietly resuming onto a different setup. The stage cache lives in its own namespace, so it can't collide with the pack's older long-video cache.

Install

ComfyUI Manager → MiniMax H3 Audio T8, or:

cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8

Full restart, then browser refresh. requirements.txt installs nothing, by design. You also need FFmpeg on PATH - the pack uses it for the crash-isolated H.264/AAC encoding and for assembling segments into the delivered MP4. Files needed: unpruned H3 base in models/diffusion_models, HyperFlow safetensors in models/hyperflow/loras, learned 3D upscaler in models/latent_upscale_models, Qwen3-VL in models/text_encoders, VAEs in models/vae.

Where it bites

Two loader owners in one graph can trip the native compiler. The pack recorded a Windows crash inside the graph compiler on the equivalent two-model split and saw it succeed on a Core started with --disable-comfy-compiler. Keep that in your back pocket before you spend an evening on it.

Sizes outside the defaults are not qualified. The node will accept other 32-grid sizes it can represent; the docs say those weren't GPU- or quality-qualified. The tested pair was 896×448 (LOW 448×224) and, in the 0.6 MP retest, 1024×576 with LOW 512×288.

"It ran" is not "it looks good". The pack's own final note for this route is that machinery-level passes are not a human review of motion, stability, dialogue or the seam - and it asks for a full watch-through, not stills.

CategoryT8/MiniMax H3/Long Video/Experimental

Inputs (27)

NameTypeDefaultDescription
model_pass1MODEL—
model_pass2MODEL—
hyperflow_fileCOMBO1 options: missing_hyperflow_weights
low_widthINT44832–16384—
low_heightINT22432–16384—
upscaler_modelCOMBO0 options:
color_matchBOOLEANtrue—
clipCLIPNative MiniMax H3 Qwen3-VL CLIP.
video_vaeVAE—
audio_vaeVAE—
chain_idSTRINGh3_hyperflow_long_video_exp—
total_duration_secondsFLOAT8.00—
widthINT89632–16384—
heightINT44832–16384—
render_window_framesINT124—
context_framesCOMBO223 options: 5, 22, 39
global_promptSTRINGUsed when Prompt Relay is disabled. With Relay, leave empty or copy the Plan global prompt exactly.
segment_prompts_jsonSTRINGPrompt overrides must be empty when Prompt Relay owns the timeline.
minimum_free_vram_mibINT5120–65536Rechecked before every segment; this is a start floor, not a peak guarantee.
base_seedINT1234567890–18446744073709550000—
seed_policyCOMBOincrement3 options: increment, fixed, hash_chain_segment
resume_existingBOOLEANtrue—
filename_prefixSTRINGH3_HyperFlow_Long_Video_EXP—
audio_seam_policyCOMBOcosine_bridge2 options: cosine_bridge, none
bridge_msFLOAT5.00–50—
bit_depthCOMBO82 options: 8, 10
crfINT180–51—

Outputs (6)

NameTypeDescription
videoVIDEO—
video_pathSTRING—
manifest_pathSTRING—
completed_segmentsINT—
statusSTRING—
report_jsonSTRING—