Nodes/comfyui-minimax-h3-audio-T8/MiniMax H3 Progressive Dual Setup / 渐进双模型配置 (EXP/T8)
ComfyUI Node

MiniMax H3 Progressive Dual Setup / 渐进双模型配置 (EXP/T8)

One clock, two passes, zero generation

By T8mars·Created 2 months ago·Updated about 7 hours ago· 1,158
MiniMax H3 Progressive Dual Setup / 渐进双模型配置 (EXP/T8)
  • model
  • model_hires
  • model
  • model_hires
  • sampler
  • sigmas
◄steps8►
◄shift_video12.0►
◄shift_audio3.0►

This node generates nothing. It's the setup half of the pack's progressive sampling route, and if you're coming from image generation it's the closest video analogue to a low-res first pass followed by a hires fix - except here the "hires fix" is a learned 3D latent upscaler, and instead of restarting the sampler you finish the same Euler schedule at full resolution.

The payoff the author measured is time, not VRAM. On fixed short clips, splitting an 8-step run into 6 low-res + 2 full-res cut end-to-end wall time by roughly 35–38% (T2VA 126.15 s → 79.03 s in the warm A/B repeat), with human review calling the picture comparable. The setup node is what makes both passes share one schedule so that math is even coherent.

What it does

It builds a tiny CPU tensor purely to check the audio/video channel layout, then clones your model twice and patches each clone's model_sampling with the video and audio sigma shifts, writing minimax_h3_sigma_shift_video and minimax_h3_sigma_shift_audio into the transformer options. It returns an Euler sampler object and a single full SIGMAS table, and it raises Progressive phase schedules differ if the two branches don't end up with identical schedules.

One thing to internalize: your input MODEL is not modified in place, and this node doesn't sample, doesn't encode a prompt, and doesn't need an empty latent from you.

The inputs and outputs

  • model - the LOW branch: base H3 → that branch's own H3-compatible LoRA chain → exactly one already-adapted attention backend.
  • steps (default 8, 2–1000) - the full native schedule length, not the number of low-res steps. The downstream sampler decides how many of those run cheap.
  • shift_video (default 12) and shift_audio (default 3) - the H3 shift values the pack standardizes on. Drift off these and audio is the first thing to sound wrong; the pack's own troubleshooting order is sampler, scheduler, steps, shifts, LoRA, in that order.
  • model_hires (optional) - the HIGH branch, wired the same way as LOW. Leave it unconnected and HIGH just reuses the LOW source config. What you can't do is sneak a second-pass LoRA in after this node and call it a separate HIGH setup - the two branches have to match in architecture, latent format, and AV clock, even when their weights differ.

Outputs: two MODELs (model and model_hires), a SAMPLER, and SIGMAS. Those four go straight into the progressive sampler or the Progressive Long Video node. If you're using TST, it goes after this node, and both TST nodes take the same full SIGMAS table - the HIGH pass uses its slice of that table and must not be handed a truncated one.

Install and what you need alongside it

Manager, search MiniMax H3 Audio T8; or cd ComfyUI/custom_nodes && git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8, then a full restart. The requirements file is empty by design - nothing to pip install.

The whole route additionally needs the learned 3D latent upscaler at ComfyUI/models/latent_upscale_models/minimax_h3_latent_upscaler_3d_fp16.safetensors, which you choose in the downstream node. It is not shipped and not auto-downloaded; it comes with the OpenVDN model pack. Why learned and not just latent resize? Because a plain resize of a video latent wrecks temporal coherence - the upscaler exists precisely so the two halves of the schedule are talking about the same content. That's the same reason step-distilled few-step pipelines are the other lever on video cost: fewer steps is less work at every stage, but only if the schedule stays coherent.

Traps

Don't mix model/sampler/sigmas from one Setup with a sampler from another - they're one contract. Don't stack accelerators on these branches: OpenVDN, SLA, VSA, Sol and BlockCache each want to own the model or attention, and this route already assigns exactly one backend per branch. And if ComfyUI's nodes all go red, update core, frontend and Manager before you start debugging the plugin.

CategoryT8/MiniMax H3/Performance/Experimental

Inputs (5)

NameTypeDefaultDescription
modelMODEL—
stepsINT82–1000—
shift_videoFLOAT12.00.01–100—
shift_audioFLOAT3.00.01–100—
model_hiresoptMODEL—

Outputs (4)

NameTypeDescription
modelMODEL—
model_hiresMODEL—
samplerSAMPLER—
sigmasSIGMAS—