H3 Progressive · Plan + LOW Source (T8 EXP)
The Node That Decides Where Your First Pass Ends
- high_source
- full_sigmas
- plan
- low_source
- low_sigmas
- high_sigmas
- report_json
Progressive sampling is the honest version of "generate small, then upscale": you run the first few denoising steps at a small canvas, hand the latent to a learned 3D latent upscaler, and finish the remaining steps at full size - all on one schedule, so the video never restarts from scratch. MiniMax H3 made it feasible because the weights and day-one ComfyUI support landed together, and because the second-resolution trick actually changes nothing about the model's time grid.
This node is the brain of that. It plans the split and produces the small source latent. That's it. No MODEL, no conditioning, no sampling, no upscaler.
What it does
Give it your full native Euler sigma table and the latent you intend to end up with, and it hands back a plan object plus:
- plan - the typed
T8_PROGRESSIVE_STAGE_PLANthat every other node in this family reads. It's the contract; save/load nodes will refuse artifacts that don't carry one. - low_source - the small-canvas source latent, derived from your high-resolution source.
- low_sigmas - the front slice of the schedule,
full_sigmas[:low_evaluations+1]. - high_sigmas - the tail,
full_sigmas[low_evaluations:]. Note the overlap at the boundary: that's the handoff point, not a bug. - report_json.
The split is a slice of one table, not two independent schedules. That's what keeps the audio clock continuous across the boundary instead of rebasing it.
Inputs that matter
- high_source - the AV latent at your target canvas. Empty for text-to-video, or the first-frame-conditioned latent for image-to-video.
- full_sigmas - the whole native H3 schedule. Don't pre-slice it; this node does the slicing and the stages need the original to verify against.
- low_evaluations (default 4) - how many model evaluations run at the small size. With eight total steps, 4 means 4+4; 6 means 6+2. Leave at least one evaluation for HIGH.
- low_scale (default 0.5) - the spatial ratio of the small canvas. It's a width/height factor, not a time or VRAM factor: 0.5 does not mean a quarter of the memory or half the duration, whatever your intuition says.
- task -
t2vaori2va. The low-resolution conditioning for a first-frame task is produced by scaling the latent, while HIGH keeps the original guide - it is not a re-encoded small version of your reference image. - input_mode -
emptyfor the normal path, or the explicit initialized-AV experimental option.
Outputs go to ONE Stage Conditioning (or the Relay variant), LOW Sampler Only, and the save/load nodes. LOW's boundary, not this plan, is what the HIGH side actually consumes.
Install
Manager, search MiniMax H3 Audio T8. Or:
cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8
Fully restart ComfyUI, refresh the browser page. Nothing in the pack's requirements.txt gets installed - that file exists to not touch your Torch/CUDA stack. You need current ComfyUI with native H3 support; the author's examples assume the H3 base model, Qwen text encoder and both VAEs are already in models/diffusion_models, models/text_encoders and models/vae.
The learned upscaler is a separate download placed in models/latent_upscale_models/ - it is not shipped with the nodes, and this plan node doesn't load it.
Where people get burned
Planning on t2va, then wiring a first frame in. The public interface covers text-to-video and single-first-frame image-to-video; multi-reference, tail frames, regional control and audio locking are not part of this route. The node's own description draws that line explicitly.
Reaching for this as a general upscaler. It isn't one - you're not upscaling a finished video, you're finishing a half-denoised latent at a bigger size, which is why the upscaler has to be the matching learned 3D latent upscaler rather than an RGB model.
And the usual EXP-level caveat, which the author states more candidly than most packs do: verified performance claims are single fixed clips on a specific card. The progressive route passing on one T2VA sample is not a promise about your LoRA stack, your duration or your GPU. If you needed a general speedup guarantee before you start, this isn't that node - but for iterating on composition at a smaller canvas, it's a genuinely good deal.
Start from examples/workflows/44-progressive-split/ and pick Full_NoSave first, so you can see the plan, the boundary and the HIGH result in one run before you start banking artifacts.
Inputs (6)
| Name | Type | Default | Description |
|---|---|---|---|
| high_source | LATENT | — | |
| full_sigmas | SIGMAS | — | |
| low_evaluations | INT | 41–999 | — |
| low_scale | FLOAT | 0.500.25–0.99 | — |
| task | COMBO | t2va | 2 options: t2va, i2va |
| input_mode | COMBO | empty | 2 options: empty, initialized_av_exp |
Outputs (5)
| Name | Type | Description |
|---|---|---|
| plan | T8_PROGRESSIVE_STAGE_PLAN | — |
| low_source | LATENT | — |
| low_sigmas | SIGMAS | — |
| high_sigmas | SIGMAS | — |
| report_json | STRING | — |