Nodes/comfyui-minimax-h3-audio-T8/H16-3 · Sample ONE PASS2 Window (T8 EXP)
ComfyUI Node

H16-3 · Sample ONE PASS2 Window (T8 EXP)

One H3 window, five things you can edit — the PASS2 node that stops hiding the sampler

By T8mars·Created 2 months ago·Updated about 7 hours ago· 1,158
H16-3 · Sample ONE PASS2 Window (T8 EXP)
  • model
  • positive
  • source_segment
  • lifted_segment
  • segment_spec
  • pass2_context
  • plan
  • noise
  • sampler
  • sigmas
  • previous_result
  • negative
  • cumulative_av_latent
  • window_result
  • core_segment_result
  • report_json
◄audio_outputpreserve_first_pass►
◄cfg1.00►

The all-in-one second-pass nodes in this pack bake in their own sampler, sigmas and noise. That's convenient right up until you want to swap the scheduler for one window, or try a different Turbo LoRA strength on the tail of a clip, and discover you can't.

This node is the opposite: it samples exactly one temporal window of the H16-3 second pass and takes model, positive, noise, sampler and sigmas from outside. Bring your own BasicGuider/SamplerCustomAdvanced chain, wire it in, done.

The shape of a window

Each window covers a slice of the timeline plus 17 frames of guarded overlap with its neighbour, and it lifts that slice through the learned 3D latent upscaler before sampling. The previous window's output is passed in as a typed window_result rather than re-derived, which is what lets the window chain stay honest: the node checks the previous result's plan SHA, source identity, window index and audio policy, and refuses to continue if any of them moved.

Because of that check, the first window is special - it must have no previous_result (a previous_result on window 0 is an error, not a warning), and after that every window needs the one immediately before it. Window 5 wants window 4's result, not window 2's.

The one input with real consequences

audio_output defaults to preserve_first_pass, and that default is the right answer for almost everybody. The first pass already produced audio; this second pass is refining the picture. preserve_first_pass keeps it, which means the failure boundary is small - if anything goes wrong with the experimental joint pass, you still have a soundtrack.

refined_exp is the other option, and it's the one that's actually experimental: on the final window it applies the old absolute-time audio crossfade and quiet-tail policy, placing each chunk's audio on absolute video-frame coordinates and crossfading only in overlaps. If the merge fails it falls back to first-pass audio rather than dropping video. Try it on a fixed seed and a fixed clip, on purpose, and compare. Don't put it on a client job.

cfg defaults to 1, which suits the distilled few-step Turbo LoRAs people run on H3. Raise it and you need a real negative conditioning plugged into the optional socket.

Inputs and outputs, quickly

Required: model, positive, source_segment, lifted_segment, segment_spec, pass2_context, plan, noise, sampler, sigmas, audio_output, cfg. Optional: previous_result and negative.

Out: cumulative_av_latent - the full accumulated AV latent so far, which is what you decode or feed to the next window - plus window_result (typed, goes to the next window, or to a save node), core_segment_result and report_json.

One trap inherited from the old node, which the author calls out in the description: the legacy denoised_output was only an output alias, not a clean x0. If you're following an old workflow that treats it as a predicted clean image, you're reading the wrong socket.

Install

Manager → MiniMax H3 Audio T8, or clone into ComfyUI/custom_nodes, then full restart:

cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8

No pip packages - the pack's requirements.txt is intentionally empty so it can't touch your Torch/CUDA install. You will need the H3 base model, Qwen3-VL encoder, both VAEs, and minimax_h3_latent_upscaler_3d_fp16.safetensors in models/latent_upscale_models/. Newer ComfyUI core required; if every node is red, update core, frontend and Manager together, then restart.

Where it bites

The doc language around this whole family is unusually honest and worth taking literally: seven-window tiny tests passing is not a claim that your 124-frame run at full resolution with your LoRAs will look the same. Plan for one window's worth of tuning before you scale out to seven.

Practically, the two failure modes are: stale previous result (you changed the plan, the model, or the audio policy between windows - the node is right to refuse, so rebuild from window 0), and forgetting that sigmas is per-window. The sigmas here are the second pass's absolute schedule; feeding them a first-pass schedule will produce a very expensive blur. And on 16GB cards, don't run two H3 jobs at once - the pack is explicit about that.

CategoryT8/MiniMax H3/Modular Sampling/H16 Experimental

Inputs (14)

NameTypeDefaultDescription
modelMODEL—
positiveCONDITIONING—
source_segmentLATENT—
lifted_segmentLATENT—
segment_specT8_CHUNKED_SOURCE_SEGMENT—
pass2_contextT8_CHUNKED_PASS2_CONTEXT—
planT8_H3_CHUNKED_TWO_PASS_PLAN—
noiseNOISE—
samplerSAMPLER—
sigmasSIGMAS—
audio_outputCOMBOpreserve_first_pass2 options: preserve_first_pass, refined_exp
cfgFLOAT1.000–100—
previous_resultoptT8_H16_PASS2_RESULT—
negativeoptCONDITIONING—

Outputs (4)

NameTypeDescription
cumulative_av_latentLATENT—
window_resultT8_H16_PASS2_RESULT—
core_segment_resultT8_CHUNKED_PASS2_RESULT—
report_jsonSTRING—