Nodes/comfyui-minimax-h3-audio-T8/H3 Native Dual · Separate LOW4 / LOW20 / HIGH3-5 (T8 EXP)
ComfyUI Node

H3 Native Dual · Separate LOW4 / LOW20 / HIGH3-5 (T8 EXP)

One stage of the dual recipe at a time, without the loop hiding under you

By T8mars·Created 2 months ago·Updated about 7 hours ago· 1,158
H3 Native Dual · Separate LOW4 / LOW20 / HIGH3-5 (T8 EXP)
  • model
  • av_latent
  • model
  • sampler
  • sigmas
  • stage_context
  • report_json
◄stagedual_low_4►
◄shift_video12.0►
◄shift_audio3.0►

What it is

The pack's dual-model long-video runner is a big all-in-one node: it plans LOW, runs it, upsamples, reconciles audio and runs HIGH inside one black box. Which is great until you want the HIGH half to use a different model, a different LoRA, a different prompt or a different noise seed - at which point you're fighting the box.

This node gives you one stage of that recipe, on its own, with no loop and no sampling. Instantiate it twice (or more) and each instance is an independently editable stage. Outputs: model, sampler, sigmas, stage_context, report_json. The report carries the legacy schedule and the planned NFE so you can see what you're about to run.

The stage table

stage is a dropdown with five entries, and they are not interchangeable:

| stage | what it is | handoff | |---|---|---| | dual_low_4 | Core simple8's first four intervals, terminal sigma non-zero | denoised_output → learned 3D upscaler; audio unfinished | | dual_low_20 | the full native_flow 20-step table | denoised_output → upscaler; audio already complete | | dual_high_3 / _4 / _5 | each stage's own published LBH raw-video sigma table | terminal output → decode |

Two things follow from that table and they matter more than anything else on this page.

HIGH is not "the rest of LOW's trajectory." It's a separate descent with its own published table and its own MODEL. The pack says this explicitly because a natural but wrong assumption is that a HIGH pass is just the second half of the schedule.

It isn't FastH3 V2's DMD schedule either. Different route, different numbers, and swapping one in gets you a result that isn't either recipe.

Inputs

  • model and av_latent - the stage's model and the AV latent it will sample.
  • stage - one of the five above.
  • shift_video / shift_audio - default 12 / 3. They must be positive and finite, and they're recorded in the stage context: the binding checks the MODEL's own clock-shift options against the context, so a stage whose model was configured differently will fail rather than run with a mismatched clock. Set these to match the MODEL you're actually handing it.

stage_context is a descriptor of a legacy stage - not tensor provenance, not a completion receipt. Ordinary ComfyUI caching still governs what re-runs in your session, and if you want a frozen stage you use the pack's save/load pair rather than hoping.

Wiring a full dual run

LOW Setup(dual_low_4 or _20) → sampler → denoised_output
   → learned 3D upscaler → Native Dual Handoff → HIGH Setup(dual_high_N) → sampler → decode

And examples/workflows/42-dual-model-split ships 24 graphs covering LOW4/LOW20 × HIGH3/4/5 in minimal, effects, save and cold-HIGH variants. Start there.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8

Or ComfyUI Manager → search MiniMax H3 Audio T8 → install → then fully exit and restart and refresh the page. If nodes go red or vanish, update ComfyUI core, the frontend and Manager together; the pack's troubleshooting notes that updating only the custom node often isn't enough, because these nodes follow native H3 APIs in core.

Models: transformer in models/diffusion_models, Qwen3-VL text encoder in models/text_encoders, video and audio VAEs in models/vae, LoRAs in models/loras - including the learned 3D upscaler and the EMA-B acceleration LoRA where the workflow names it. Don't swap acceleration LoRAs casually: the README states plainly that generic EMA, Ref2VA and OpenVDN turbo LoRAs are not interchangeable. The repo ships no weights, and requirements.txt installs no packages on purpose.

Common issues

"Clock shifts differ from context." You built the stage with one shift pair and handed it a MODEL configured with another. Set the shifts to match - this check exists precisely because a silent clock mismatch produces audio drift that looks like a model problem.

The audio is wrong after the handoff. That's the Handoff node's policy choice, not this one: 4 + auto continues unfinished audio, 20 + auto preserves completed audio.

You used LOW's output for the upscale. Use denoised_output. It's the documented handoff for the LOW stages in this route.

Save/resume of a LOW stage fails. You need real path and SHA values, a sampler the pack certifies for the stage, and a recovery graph that genuinely deletes the LOW model/conditioning/sampling chain - the shipped cold-HIGH graphs are built to do that, so start from one.

CategoryT8/MiniMax H3/Modular Sampling/Experimental

Inputs (5)

NameTypeDefaultDescription
modelMODEL—
av_latentLATENT—
stageCOMBOdual_low_45 options: dual_low_4, dual_low_20, dual_high_3, dual_high_4, dual_high_5
shift_videoFLOAT12.00.01–100—
shift_audioFLOAT3.00.01–100—

Outputs (5)

NameTypeDescription
modelMODEL—
samplerSAMPLER—
sigmasSIGMAS—
stage_contextT8_STAGE_CONTEXT—
report_jsonSTRING—