H3 SPEED · DCT + AV Transition (T8 EXP)
The DCT trick that hands one SPEED stage to the next (and its own audio clock)
- completed_stage
- next_stage
- noise
- report_json
Two inputs, one output. MiniMaxH3SPEEDDCTTransitionEXPT8 takes a finished SPEED stage and the next stage's freshly prepared canvas, and returns the exact NOISE that the next stage should start from. No sampling, no model call, no silent reuse of an old stage.
If you're not on the SPEED split route, this node is irrelevant and you can go back to your life. If you are, it's the only thing standing between you and a broken picture.
What it's doing, really
SPEED is the research idea that you generate a clip at growing spatial resolution, stage by stage: start small, then expand. Resolution changes can't just be resized away in latent space - you have to move the existing low-frequency content into the right corner of a bigger spectrum and fill the rest with appropriately-shaped noise. That's what DCT expansion does here: the pack implements the official DCT-II basis, expands the completed stage's latent coefficients into the next stage's larger grid, and - the part people forget - reindexes the audio.
H3's audio is generated jointly. When you change the video canvas between stages, the audio state has to be re-anchored to the new flow position rather than left where it was, otherwise you get drift, DC offset or the dreaded spectral mush. The node's own description is honest about what it doesn't do: "Does not sample or silently reuse old stages." It solves the next segment's noise and reports what it did in report_json.
Inputs:
completed_stage- aT8_SPEED_STAGE_RESULT, i.e. the output ofMiniMaxH3SPEEDStageSaveEXPT8orMiniMaxH3SPEEDStageSampleEXPT8. Note it's the public-flow state the sampler produced, not a decoded x0.next_stage- aT8_SPEED_STAGE_SPECfromMiniMaxH3SPEEDStageSetupEXPT8for the stage you're about to run. The two have to agree: a mismatched pair is a hard error, not a warning.dct_chunk_size(default 64, 1–1024) - the workspace knob. Bigger chunks mean fewer, larger DCT work buffers; if you're tight on VRAM, drop it.
Output: noise (wire it into the next stage's MiniMaxH3SPEEDStageSampleEXPT8) and report_json.
The honest read on SPEED
Before you spend an evening on this: the pack's own measurement notes say the fixed SPEED plan lost to the baseline in a same-input Stock20 run (≈248.7s / 16.2 GiB vs 243.2s / 12.5 GiB), and their blind review picked the baseline. So the accelerated path that passed every numeric contract still didn't win. The split-workshop example docs literally say do not read these values as a quality or speed recommendation.
That doesn't make the node useless - it's how you reproduce the research and inspect it stage by stage instead of through one opaque whole-chain sampler - but it does mean you should treat "SPEED" here as a lab name, not a promise.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8
Exit ComfyUI fully and restart, then hard-refresh the frontend. Manager search term: MiniMax H3 Audio T8. The pack ships an intentionally empty requirements.txt - no extra pip packages - and depends on a recent ComfyUI core for the native H3 model class. Models go in models/diffusion_models, the Qwen3-VL encoder in models/text_encoders, both VAEs in models/vae.
Gotchas
The 50-speed-split example set (82 graphs) is the reference wiring; start from none_none_full_save before you look at the effect-laden ones. Cold-resume graphs select a saved artefact by relative manifest path and exact SHA - the placeholders in the JSON are not runnable, and an edited past stage will not be auto-recomputed for you. Bounding your dct_chunk_size is the first lever if the transition itself is what's spiking memory; if you're on 16 GB total, run one H3 job at a time and don't try to expand into a big canvas and a big frame count at once.
Inputs (3)
| Name | Type | Default | Description |
|---|---|---|---|
| completed_stage | T8_SPEED_STAGE_RESULT | — | |
| next_stage | T8_SPEED_STAGE_SPEC | — | |
| dct_chunk_size | INT | 641–1024 | — |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| noise | NOISE | — |
| report_json | STRING | — |