Nodes/comfyui-minimax-h3-audio-T8/MiniMax H3 Progressive Long Video / 渐进双采长视频 (EXP/T8)
ComfyUI Node

MiniMax H3 Progressive Long Video / 渐进双采长视频 (EXP/T8)

An entire long-video pipeline hiding in one node

By T8mars·Created 2 months ago·Updated about 7 hours ago· 1,158
MiniMax H3 Progressive Long Video / 渐进双采长视频 (EXP/T8)
  • model
  • clip
  • video_vae
  • audio_vae
  • sampler
  • sigmas
  • model_hires
  • prompt_relay_plan
  • first_frame
  • drive_audio
  • final_audio
  • video
  • video_path
  • report_json
◄chain_idh3_progressive_long_video►
◄total_duration_seconds8.00►
◄width896►
◄height448►
◄render_window_frames124►
◄context_frames22►
◄global_prompt►
◄segment_prompts_json►
◄upscaler_model▾►
◄base_seed2609032101►
◄seed_policyincrement►
◄low_evaluations4►
◄low_scale0.50►
◄precisionfp16►
◄guide_resizelegacy_bilinear►
◄context_audiovideo_and_audio►
◄audio_modenative►
◄audio_denoise_strength0.35►
◄resume_existingtrue►
◄filename_prefixH3_Progressive_UNREVIEWED►
◄audio_seam_policycosine_bridge►
◄audio_bridge_ms5►
◄crf18►
◄query_chunk_rows256►
◄reserve_vram_mib2048►
◄eav_modedisabled►
◄eav_tau4.0►
◄eav_start_video_progress0.15►
◄eav_end_video_progress0.90►
◄eav_max_workspace_mib32►
◄eav_g_hard_limit1.50►

Every long-video workflow in ComfyUI eventually becomes the same puzzle: generate a window, stitch it to the last one, keep the audio from drifting, and don't lose an hour of work when a segment fails. This node does all of that internally. You give it a configured model pair, a sampler, one full sigma table, CLIP and the two VAEs, plus a duration - it renders the video in segments, splits each segment's schedule into a cheap low-res first pass and a full-res finish, carries the accepted picture forward as context, and writes the MP4 itself.

It is an output node. You don't add a SaveVideo afterwards.

Defaults that tell you the author's mood

filename_prefix defaults to H3_Progressive_UNREVIEWED, and the code raises an error if the delivery report ever claims human qualification. That is not decoration - it's the pack's way of saying this route has real mechanical verification (two 8-second segments with real weights, strict video and audio decode, human review) but that a full combination matrix and quality sign-off are still open. Take it seriously. Run it, watch it, decide for yourself.

The inputs worth touching

Wire in model, clip, video_vae, audio_vae, sampler and sigmas from the Progressive Dual Setup node - those six define the run and they must come from the same contract.

  • chain_id - the identity of this run. Same id plus identical inputs means "verify and resume"; change the model, LoRA, prompt, source material, duration or resolution and you should change the id rather than drag a failed segment's cache into a new experiment.
  • total_duration_seconds, width, height - target output, dimensions snapped to 32.
  • render_window_frames (124–362, on a 17n+5 grid) - the internal generation window. Here's the counterintuitive bit: even a 3-second output renders at least a 124-frame window and trims down. You cannot ask it for 73.
  • low_evaluations - how many steps of the full 8-step table run cheap. 4 gives you 4+4, 6 gives you 6+2. The node requires at least one HIGH step.
  • low_scale - the LOW spatial fraction. At 0.5 with an 896×448 target, the first pass is 448×224.
  • context_frames - 22 or 39 frames of accepted-picture context carried into the next segment.
  • upscaler_model - pick your existing learned 3D latent upscaler. Nothing is auto-downloaded.
  • global_prompt / segment_prompts_json - the standing description of the scene versus per-segment events. One-off dialogue belongs in the segment events, not the global prompt you resend every window.
  • resume_existing - true re-verifies real models, sources, settings and on-disk receipts before reusing stages. false refuses to touch an existing chain inside an OS lock and never deletes your files.
  • audio_mode, drive_audio, final_audio, audio_seam_policy (a 5 ms cosine bridge by default) - the audio side. final_audio is preferred as the output track; being the output track doesn't mean it drove generation.
  • first_frame - a single first frame for segment 0. That's the extent of the image conditioning.
  • model_hires (optional) plus the optional prompt_relay_plan, and the eav_* group, which applies the pack's EAV gain internally. Don't feed in a MODEL that already carries an external EAV wrapper, and don't run the TST node and the internal EAV settings as two competing configurations of the same branch.

Outputs: video (a VIDEO object), video_path (the rendered file), and report_json (per-stage record - read it when a segment decides to redo itself).

Timing math you'll otherwise get wrong

An 8-second, 24 fps output is 192 frames, and with a 124-frame window the seam lands around frame 124, about 5.17 seconds in. It is not two 4-second halves. Short clips pass; that says nothing about a 24- or 32-second run.

Scope, and the honest gaps

T2VA and single-first-frame I2VA only. Not multi-reference, not end-frame, not reference-audio editing. For those, use the pack's older dual-MODEL in-node long-video route - this node doesn't replace it, and the old workflow's defaults didn't change.

OOM is still OOM: lower resolution, frames and reference count, run one H3 job at a time, and remember reserve_vram_mib checks the boundary rather than capping your card. Audio weirdness gets the same checklist as everywhere else in this pack - sampler, scheduler, steps, video/audio shifts, LoRA, before you blame the node. And installing is the usual: Manager → MiniMax H3 Audio T8 → full restart, no pip extras, plus ffmpeg on PATH so the delivery step can write the file.

CategoryT8/MiniMax H3/Long Video/Experimental

Inputs (42)

NameTypeDefaultDescription
modelMODEL—
clipCLIP—
video_vaeVAE—
audio_vaeVAE—
samplerSAMPLER—
sigmasSIGMAS—
chain_idSTRINGh3_progressive_long_video—
total_duration_secondsFLOAT8.00—
widthINT89664–8192—
heightINT44864–8192—
render_window_framesINT124内部窗口至少124帧、17n+5网格,无固定帧数上限;短片按总时长裁出。窗口越大,显存和耗时通常越高。
context_framesCOMBO222 options: 22, 39
global_promptSTRING—
segment_prompts_jsonSTRING—
upscaler_modelCOMBO0 options:
base_seedINT26090321010–18446744073709550000—
seed_policyCOMBOincrement2 options: increment, fixed
low_evaluationsINT41–999完整8步表中填4为4+4,填6为6+2;至少留1步HIGH。
low_scaleFLOAT0.500.25–0.99—
precisionCOMBOfp163 options: fp16, bf16, fp32
guide_resizeCOMBOlegacy_bilinear2 options: legacy_bilinear, preserve_mean
context_audioCOMBOvideo_and_audio2 options: video_and_audio, video_only
audio_modeCOMBOnative3 options: native, lock_source, remix_source
audio_denoise_strengthFLOAT0.350–1—
resume_existingBOOLEANtrue相同输入重新执行会核验后复用阶段/成片。关闭则拒绝已有数据链;不清理旧结果。
filename_prefixSTRINGH3_Progressive_UNREVIEWED—
audio_seam_policyCOMBOcosine_bridge2 options: none, cosine_bridge
audio_bridge_msFLOAT50–50—
crfINT180–51—
query_chunk_rowsINT25632–2048—
reserve_vram_mibINT2048512–32768—
eav_modeCOMBOdisabled3 options: disabled, report_only, apply_exp
eav_tauFLOAT4.0-32–32—
eav_start_video_progressFLOAT0.150–1—
eav_end_video_progressFLOAT0.900–1—
eav_max_workspace_mibINT324–512—
eav_g_hard_limitFLOAT1.501–3—
model_hiresoptMODEL—
prompt_relay_planoptH3_T8_PROMPT_RELAY_PLAN—
first_frameoptIMAGE—
drive_audiooptAUDIO—
final_audiooptAUDIO—

Outputs (3)

NameTypeDescription
videoVIDEO—
video_pathSTRING—
report_jsonSTRING—