H16-3 分块 PASS2 · 音频精修(可选)
The H3 second-pass sampler that quietly fixes your audio
- noise
- guider
- sampler
- sigmas
- latent_image
- output
- denoised_output
The display name says "audio refinement", which undersells it and slightly misleads: DeciiaChunkedPass2Sampler is a second-pass sampler. It does the H3 low-to-high refinement in temporal chunks so a long clip fits in VRAM, and the audio policy is one dropdown on top of that. If you have a 20-second H3 clip and PASS 2 is either OOMing or producing seams, this is the node.
Why PASS 2 needs chunking at all
MiniMax H3 generates picture and sound together in one latent, so anything that re-samples the latent is re-sampling the audio too. A one-shot PASS 2 over a long window is exactly the thing that eats a 16 GB card. The usual answer is temporal chunking - process 34 frames, step forward, blend - and the usual failure mode is that the joins are visible and audible.
This node is the T8 pack's wrapper around its own formal v4 chunked two-pass executor. It isn't a third-party reimplementation: the node ID was kept compatible with workflows imported from the Deciia lineage (issue #18), but the executor underneath is T8's own, and no external code or weights are bundled.
How it actually works
It reads the video and audio halves out of the nested H3 AV latent, pulls the raw conditioning, CFG and model out of the guider you hand it, and builds a chunked plan. Spatial strategy is fixed to full-frame - there is no hidden spatial tiling, so one chunk is always the whole frame. Only time is split.
guarded_overlap_exp chunks with overlap and blends the overlap; full_clip_safe turns chunking off and does the whole clip in one go, which is what you want only if it fits. The two integer knobs move in 17-frame steps because that's H3's temporal compression grid - 34 frames of chunk with 17 of overlap is the default, i.e. 17 new frames per step.
Then the audio decision:
preserve_first_pass(the default) keeps the audio your first pass already made. Minimum blast radius, and the honest recommendation for a first run.refined_exppublishes the audio captured during the chunked second pass, placed back on the absolute audio timeline, crossfaded only where chunks overlap, with quiet tails protected by first-pass audio. If the merge fails, it falls back to first-pass audio rather than dropping your video.
Inputs and outputs worth wiring
Everything upstream of this is your normal PASS 2 chain: noise, guider, sampler, sigmas, and latent_image (a native nested H3 AV latent - the node rejects anything else, including a plain video latent).
temporal_chunk_frames / temporal_overlap_frames are the only knobs most people touch. anchor_strength (0.999) controls how hard the first pass anchors the overlap region - drop it and you'll see the join move. audio_output picks the two policies above.
Outputs are output (a LATENT - decode this one) and denoised_output, which the tooltip is upfront about: it's a placeholder equal to output, kept so the node plugs into sampler sockets that expect that name. Decoding both gives you two identical videos and wasted time.
One detail you can't see in the schema: the chunked plan uses minimax_h3_latent_upscaler_3d_fp16.safetensors as its learned upscaler, so that file needs to be in models/latent_upscale_models.
Installing the pack
ComfyUI Manager → search MiniMax H3 Audio T8, or by hand:
cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8
Fully quit ComfyUI and start it again, then refresh the browser - restarting only the frontend leaves the new Python unloaded. requirements.txt installs no packages on purpose: the base nodes use the torch/numpy/Pillow/safetensors ComfyUI already has, and optional features check their own dependencies when you use them. You need a recent ComfyUI with native H3 support, the H3 base in models/diffusion_models, Qwen3-VL in models/text_encoders, and the video/audio VAEs in models/vae.
Where people get burned
The latent has to be the real thing. The node checks shapes ([1,24,...] video, [1,32,2,...] audio) and finiteness before doing anything. If you feed it the output of a node that flattened the nested AV latent, you get a shape error that has nothing to do with chunk sizing.
Chunk values off the 17-grid won't do what you think. Stick to multiples of 17.
Don't expect refined_exp to be a free win. It's the experimental path, qualified on one specific 73-frame GPU sample. The safe order is: run preserve_first_pass, confirm the picture, then re-run with the same seed and chunk settings and A/B the audio. And remember the fallback is deliberate - a silent return to first-pass audio is a real possibility, not a bug.
Inputs (10)
| Name | Type | Default | Description |
|---|---|---|---|
| noise | NOISE | — | |
| guider | GUIDER | — | |
| sampler | SAMPLER | — | |
| sigmas | SIGMAS | — | |
| latent_image | LATENT | — | |
| temporal_strategy | COMBO | guarded_overlap_exp | 2 options: guarded_overlap_exp, full_clip_safe |
| temporal_chunk_frames | INT | 3417–3600 | — |
| temporal_overlap_frames | INT | 170–1700 | — |
| anchor_strength | FLOAT | 0.9990–1 | — |
| audio_output | COMBO | preserve_first_pass | 2 options: refined_exp, preserve_first_pass |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| output | LATENT | — |
| denoised_output | LATENT | Placeholder equal to output; decode output. |