H3 Set Audio Prefix Noise Mask
Stop H3 re-recording your intro
- audio_latent
- trim_info
- audio_latent
Chaining H3 segments is how you get past one generation's worth of runtime, and the hard part in MiniMax H3 is that audio is welded to the picture. When you glue the last 22 frames of the previous segment onto the front of the next one, you want the model to see that audio - to know a voice, a song, or a room tone is already in progress - without generating it all over again. Video has a mask for this. Audio does not, unless you add one.
That's this node's entire job: attach a noise_mask to the standalone H3 audio latent - zero where the audio should survive, one where H3 should generate, the same convention the pack's Set Video Latent Noise Mask uses for video.
Why it needs its own node
Until you run H3 Concat AV Latent, H3's video and audio live as separate latents: video in [B,C,T,H,W], audio in [B,32,2,T]. A video noise mask says nothing at all about the audio stream, so a properly masked video continuation with an unmasked audio latent will happily resynthesise speech over the top of dialogue you already had.
The other half of the job is the boundary. Video Continuation Concat records how many waveform samples belong to the prefix; the mask needs a count of audio latent positions. This node converts between them, rounding in favour of protecting slightly too much rather than handing a sliver of your real audio to the sampler.
The inputs that matter
audio_latent is required and must be the standalone [B,32,2,T] H3 audio latent. Feed it the nested AV latent and it stops with Expected standalone H3 audio latent [B,32,2,T], got [...] - split first with H3 Separate AV Latent, mask the video with Set Video Latent Noise Mask, then join both again after this node.
mode has three settings. protect_prefix_generate_body (default) zeroes the prefix positions and sets one after the boundary. protect_all zeroes the whole latent when the audio is already finished. generate_all sets it all to one, for when there's no useful prefix audio and H3 should write the whole track.
trim_info is the token that came out of Video Continuation Concat. It's an optional socket in the schema but it is not optional in practice for the default mode: without it, prefix and total sample counts are unknown and the node raises rather than guessing. One subtlety worth knowing - if the concat node recorded that no prefix audio was supplied at all, the boundary is deliberately zero, so protect_prefix_generate_body degenerates to "generate everything". That's correct behaviour (there is no waveform to protect) and a common source of "my mask did nothing" confusion.
Output is a single audio_latent with noise_mask attached, and it goes straight into H3 Concat AV Latent alongside the masked video latent.
Install
The pack is a single clone:
cd ComfyUI/custom_nodes
git clone https://github.com/wjie98/comfyui-svdint4.git
# restart ComfyUI
Heads-up on the name: the repo has been renamed at least once. GitHub and ComfyUI Manager list it as comfyui-svdint4, while the README calls the project "ComfyUI Turing Utils" and tells you to clone comfyui-turing-utils. Same pack either way; the folder name is irrelevant to ComfyUI.
The pack's headline feature is a hand-written CUDA kernel (ConvRot quantisation, bundled W8A8/Sage/Sol attention) built with python -m pip install -v --no-build-isolation -e ./kernel. You do not need any of that for this node. It's pure Python and torch, the pack's requirements.txt adds nothing but safetensors, and ComfyUI already ships PyAV and torchaudio. Skip the CUDA build unless you're also using the loader and attention nodes.
Fair warning: I searched the community corpus for this pack, the author, and both repo names and got zero hits, so there's no forum thread to fall back on. The README is your documentation.
Things that trip people up
The usual failure is a shapes error thrown at you rather than a quiet wrong result, which is genuinely helpful here: [B,32,2,T] is checked, trim_info provenance is checked, and the sample counts are validated against each other. If you get trim_info contains an invalid audio boundary, you've wired metadata from somewhere that isn't Video Continuation Concat - that custom type only ever comes out of that one node.
Second, remember this node does not touch the samples. It only replaces/adds noise_mask. If your continuation still sounds like a fresh take on the intro, check the mask landed on the latent that actually reaches the sampler - running this after H3 Concat AV Latent won't help, because the nested AV latent is a different shape.
And third: this is the audio half of a pair. Route the same trim_info to Trim Video Continuation Prefix. Two consumers, one boundary token - hand-roll either count and the audio drifts against the picture by a few milliseconds per segment, which you discover at the end of a forty-minute render.
Inputs (3)
| Name | Type | Default | Description |
|---|---|---|---|
| audio_latent | LATENT | — | |
| mode | COMBO | protect_prefix_generate_body | 3 options: protect_prefix_generate_body, protect_all, generate_all |
| trim_infoopt | TURING_UTILS_VIDEO_TRIM_INFO | — |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| audio_latent | LATENT | — |