H3 Native Dual · Learned / Audio Handoff (T8 EXP)
The audio
- learned_latent
- highres_template
- positive
- av_latent
- positive
- report_json
Why this node has to exist
The dual-model long-video recipe goes LOW → learned 3D upscaler → HIGH. LOW runs at low resolution (4 or 20 steps), the upscaler lifts the latents to the HIGH resolution, and HIGH runs its own short descent on top with a published LBH table.
Here's the problem: H3 is a joint AV model. Audio and video share one latent and one clock. The upscaler is a video-side tool. It does not know or care what happened to the audio, so something has to reconcile the audio before the HIGH pass can run without producing a video whose soundtrack is out of step with the picture.
MiniMaxH3NativeDualHandoffEXPT8 is that something - and it's deliberately just that. It calls the original long-video runner's reconciliation path, no sampler, zero diffusion evaluations, no upscaling. If your graph gets slower when you add this node, something else is wrong.
Inputs
learned_latent- the upscaler's output.highres_template- the HIGH-resolution AV template. This is what gives the reconciliation the right geometry for the audio rows and masks; it is not the LOW result.positive- the HIGH conditioning, which gets rebuilt at the new geometry and returned.first_pass_steps-"4"or"20". Anything else raises; it isn't a free-form number.second_audio_source-auto,legacy_policy,first_pass, orhighres_template.autois the one you want unless you're deliberately reproducing old behaviour.second_audio_strength-0to1, default0.
Outputs: av_latent, positive, report_json.
What auto does
This is the heart of the node, and it's why the first_pass_steps string isn't cosmetic:
- 4 steps + auto → pass one's audio is unfinished (the schedule's terminal sigma is non-zero). So the audio keeps evolving through the handoff instead of being frozen. The template's locked regions - where the mask says "this audio is already decided" - keep the template audio; everything else carries the coarse audio forward.
- 20 steps + auto → pass one completed its audio, so the completed audio is preserved as-is.
The legacy runner also had a compatibility migration where LOW4's first_pass/0 setting gets reinterpreted as joint continuation; that still happens here and the report records it rather than hiding it.
The general shape to remember: unfinished audio should not be locked, finished audio should not be re-invented. Getting that backwards is the classic cause of a clip where the dialogue sounds like it's being chewed or where the audio suddenly changes character at the handoff.
Wiring
LOW Stage Setup → sampler → denoised_output → learned 3D upscaler
→ Native Dual Handoff → HIGH Stage Setup → sampler → decode
HIGH conditions at the upscaler's ACTUAL output geometry ─→ Handoff (positive)
The trap: LOW's denoised_output, not its terminal output. In every upscale-handoff route in this pack, denoised_output is the prediction handed forward, and output is an unfinished x_sigma state that looks superficially similar. And note the HIGH conditioning must be built at the geometry the upscaler actually produced, not a rounded resolution you decided on in advance.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8
ComfyUI Manager: search MiniMax H3 Audio T8, install, then fully exit and restart ComfyUI and refresh the page. Registry publishing and GitHub releases are separate - clone if Manager is behind. Red or missing nodes usually mean core, frontend and Manager need updating together, not just the pack.
Models: H3 transformer in models/diffusion_models, Qwen3-VL text encoder in models/text_encoders, H3 video and audio VAEs in models/vae, LoRAs in models/loras - including the learned 3D upscaler asset this handoff sits behind. The repo ships no weights and installs no Python packages (requirements.txt is intentionally bare to protect your torch/CUDA build).
Start from examples/workflows/42-dual-model-split, which has LOW4/LOW20 × HIGH3/4/5 in minimal, effects, save and cold-HIGH forms.
Common issues
"Requires first_pass_steps 4 or 20." You typed something else, or a workflow passed a step count through. It's an enum on purpose.
Audio changes character at the handoff. You're on the wrong audio source policy for the LOW length, or you used legacy_policy while expecting today's defaults. Check the report - it states the policy that was actually applied.
Video looks fine, HIGH is soft. Not this node's department: this is the audio/geometry handoff. Look at the UPS table selection and the HIGH stage's own steps.
OOM. Lower resolution/frames references, and on a 16 GB card run one H3 job at a time. Dual-model long video is one of the routes where "just try again" is the worst idea available.
Inputs (6)
| Name | Type | Default | Description |
|---|---|---|---|
| learned_latent | LATENT | — | |
| highres_template | LATENT | — | |
| positive | CONDITIONING | — | |
| first_pass_steps | COMBO | 4 | 2 options: 4, 20 |
| second_audio_source | COMBO | auto | 4 options: auto, legacy_policy, first_pass, highres_template |
| second_audio_strength | FLOAT | 0.000–1 | — |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| av_latent | LATENT | — |
| positive | CONDITIONING | — |
| report_json | STRING | — |