MiniMax H3 Progressive Long Video / 渐进双采长视频 (EXP/T8)
An entire long-video pipeline hiding in one node
- model
- clip
- video_vae
- audio_vae
- sampler
- sigmas
- model_hires
- prompt_relay_plan
- first_frame
- drive_audio
- final_audio
- video
- video_path
- report_json
Every long-video workflow in ComfyUI eventually becomes the same puzzle: generate a window, stitch it to the last one, keep the audio from drifting, and don't lose an hour of work when a segment fails. This node does all of that internally. You give it a configured model pair, a sampler, one full sigma table, CLIP and the two VAEs, plus a duration - it renders the video in segments, splits each segment's schedule into a cheap low-res first pass and a full-res finish, carries the accepted picture forward as context, and writes the MP4 itself.
It is an output node. You don't add a SaveVideo afterwards.
Defaults that tell you the author's mood
filename_prefix defaults to H3_Progressive_UNREVIEWED, and the code raises an error if the delivery report ever claims human qualification. That is not decoration - it's the pack's way of saying this route has real mechanical verification (two 8-second segments with real weights, strict video and audio decode, human review) but that a full combination matrix and quality sign-off are still open. Take it seriously. Run it, watch it, decide for yourself.
The inputs worth touching
Wire in model, clip, video_vae, audio_vae, sampler and sigmas from the Progressive Dual Setup node - those six define the run and they must come from the same contract.
chain_id- the identity of this run. Same id plus identical inputs means "verify and resume"; change the model, LoRA, prompt, source material, duration or resolution and you should change the id rather than drag a failed segment's cache into a new experiment.total_duration_seconds,width,height- target output, dimensions snapped to 32.render_window_frames(124–362, on a 17n+5 grid) - the internal generation window. Here's the counterintuitive bit: even a 3-second output renders at least a 124-frame window and trims down. You cannot ask it for 73.low_evaluations- how many steps of the full 8-step table run cheap.4gives you 4+4,6gives you 6+2. The node requires at least one HIGH step.low_scale- the LOW spatial fraction. At 0.5 with an 896×448 target, the first pass is 448×224.context_frames- 22 or 39 frames of accepted-picture context carried into the next segment.upscaler_model- pick your existing learned 3D latent upscaler. Nothing is auto-downloaded.global_prompt/segment_prompts_json- the standing description of the scene versus per-segment events. One-off dialogue belongs in the segment events, not the global prompt you resend every window.resume_existing-truere-verifies real models, sources, settings and on-disk receipts before reusing stages.falserefuses to touch an existing chain inside an OS lock and never deletes your files.audio_mode,drive_audio,final_audio,audio_seam_policy(a 5 ms cosine bridge by default) - the audio side.final_audiois preferred as the output track; being the output track doesn't mean it drove generation.first_frame- a single first frame for segment 0. That's the extent of the image conditioning.model_hires(optional) plus the optionalprompt_relay_plan, and theeav_*group, which applies the pack's EAV gain internally. Don't feed in a MODEL that already carries an external EAV wrapper, and don't run the TST node and the internal EAV settings as two competing configurations of the same branch.
Outputs: video (a VIDEO object), video_path (the rendered file), and report_json (per-stage record - read it when a segment decides to redo itself).
Timing math you'll otherwise get wrong
An 8-second, 24 fps output is 192 frames, and with a 124-frame window the seam lands around frame 124, about 5.17 seconds in. It is not two 4-second halves. Short clips pass; that says nothing about a 24- or 32-second run.
Scope, and the honest gaps
T2VA and single-first-frame I2VA only. Not multi-reference, not end-frame, not reference-audio editing. For those, use the pack's older dual-MODEL in-node long-video route - this node doesn't replace it, and the old workflow's defaults didn't change.
OOM is still OOM: lower resolution, frames and reference count, run one H3 job at a time, and remember reserve_vram_mib checks the boundary rather than capping your card. Audio weirdness gets the same checklist as everywhere else in this pack - sampler, scheduler, steps, video/audio shifts, LoRA, before you blame the node. And installing is the usual: Manager → MiniMax H3 Audio T8 → full restart, no pip extras, plus ffmpeg on PATH so the delivery step can write the file.
Inputs (42)
| Name | Type | Default | Description |
|---|---|---|---|
| model | MODEL | — | |
| clip | CLIP | — | |
| video_vae | VAE | — | |
| audio_vae | VAE | — | |
| sampler | SAMPLER | — | |
| sigmas | SIGMAS | — | |
| chain_id | STRING | h3_progressive_long_video | — |
| total_duration_seconds | FLOAT | 8.00 | — |
| width | INT | 89664–8192 | — |
| height | INT | 44864–8192 | — |
| render_window_frames | INT | 124 | 内部窗口至少124帧、17n+5网格,无固定帧数上限;短片按总时长裁出。窗口越大,显存和耗时通常越高。 |
| context_frames | COMBO | 22 | 2 options: 22, 39 |
| global_prompt | STRING | — | |
| segment_prompts_json | STRING | — | |
| upscaler_model | COMBO | 0 options: | |
| base_seed | INT | 26090321010–18446744073709550000 | — |
| seed_policy | COMBO | increment | 2 options: increment, fixed |
| low_evaluations | INT | 41–999 | 完整8步表中填4为4+4,填6为6+2;至少留1步HIGH。 |
| low_scale | FLOAT | 0.500.25–0.99 | — |
| precision | COMBO | fp16 | 3 options: fp16, bf16, fp32 |
| guide_resize | COMBO | legacy_bilinear | 2 options: legacy_bilinear, preserve_mean |
| context_audio | COMBO | video_and_audio | 2 options: video_and_audio, video_only |
| audio_mode | COMBO | native | 3 options: native, lock_source, remix_source |
| audio_denoise_strength | FLOAT | 0.350–1 | — |
| resume_existing | BOOLEAN | true | 相同输入重新执行会核验后复用阶段/成片。关闭则拒绝已有数据链;不清理旧结果。 |
| filename_prefix | STRING | H3_Progressive_UNREVIEWED | — |
| audio_seam_policy | COMBO | cosine_bridge | 2 options: none, cosine_bridge |
| audio_bridge_ms | FLOAT | 50–50 | — |
| crf | INT | 180–51 | — |
| query_chunk_rows | INT | 25632–2048 | — |
| reserve_vram_mib | INT | 2048512–32768 | — |
| eav_mode | COMBO | disabled | 3 options: disabled, report_only, apply_exp |
| eav_tau | FLOAT | 4.0-32–32 | — |
| eav_start_video_progress | FLOAT | 0.150–1 | — |
| eav_end_video_progress | FLOAT | 0.900–1 | — |
| eav_max_workspace_mib | INT | 324–512 | — |
| eav_g_hard_limit | FLOAT | 1.501–3 | — |
| model_hiresopt | MODEL | — | |
| prompt_relay_planopt | H3_T8_PROMPT_RELAY_PLAN | — | |
| first_frameopt | IMAGE | — | |
| drive_audioopt | AUDIO | — | |
| final_audioopt | AUDIO | — |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| video | VIDEO | — |
| video_path | STRING | — |
| report_json | STRING | — |