T8 末端音频静音+淡入(VIDEO EXP)
Fix the Sound, Leave the Video Alone
- video
- video
- report_json
What it is
Audio is still the thinnest layer of the local stack - the tools got bolted on after video got good enough to want a soundtrack, and they live at the edges rather than in the middle. This node is exactly that kind of edge tool: a five-second utility for one specific annoyance, namely a clip whose first fraction of a second pops, thumps or starts on a hard transient.
It takes a VIDEO, zeroes the audio for the first N frames, then eases the sound back in with a half-cosine fade. Duration, sample rate, channel layout and total sample count are preserved. It doesn't resample, doesn't shorten, doesn't denoise or normalise the rest of the track, and it doesn't touch a single video byte.
How it works
There are two code paths, and which one you get depends on what's on the wire.
A tensor VIDEO (the native components object a generation node hands you) is the clean case. The frame data passes through untouched; only the audio component is rebuilt, using that component's own real frame rate to convert "N frames" into a sample count. Nothing is approximated.
A file VIDEO gets rewritten in place, sort of. A worker process re-encodes the audio stream and copies the video stream rather than re-rendering it. If the file is a plain native MP4 with no active trim or crop, the original video bytes are what come out. If it's something else - or if Core has an active trim/crop view on it - Core's existing streaming exporter materialises that view into a temp MP4 first, which may transcode it. That work happens as a stream; it does not load the whole clip's RGB frames into memory. On Windows the codec worker sits inside the job object, so cancelling the run doesn't leave an orphan process behind.
The envelope itself is deliberately tiny: a cosine ramp whose single-sample case is treated as zero, because continuity at the mute boundary wins over a decorative one-sample ramp.
The inputs that matter
mute_first_frames(default 1) - how many leading frames' worth of audio gets zeroed. This is the number that eats your dialogue if you push it.fade_in_ms(default 10) - the half-cosine fade length after the mute.N=0alone gives you fade-only;N=0withfade_in_ms=0is a straight identity pass.enabled(default true here, but off in the pack's example workflow) - bypass returns the original object untouched.
Outputs are video (wire it on into your normal Save Video) and report_json, which is more informative than you'd expect for a fade: it reports the sample rate, total samples, mute and fade sample counts, and the exact mute/fade end times as fractions, plus a warning that long settings can suppress the first spoken word. It also reports when a video had no audio stream at all - that's a passthrough with no_audio_passthrough, not an error.
Installing
cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8
Restart fully, refresh the browser, or install through Manager by searching MiniMax H3 Audio T8. No model downloads for this one, and the pack's requirements.txt is intentionally empty - the container library it uses is the one ComfyUI already ships, so nothing here will reinstall your Torch stack.
Where people get burned
Start at the defaults: 1 frame, 10 ms. That's the author's own advice, and it's right. Push mute_first_frames to something like 5 and you don't remove a click, you remove the first syllable. If the defaults don't fix it, a longer envelope probably isn't the answer.
If there's no click, leave it off. The pack's diagnostics example ships with this branch disabled. Enabling it just in case is how you end up quietly clipping the start of a good take.
It's a cosmetic fix, not a diagnosis. Trimming the head of the track hides an abrupt onset; it doesn't tell you why one was there. If you're trying to find the cause, look upstream in the generated or assembled audio, not here.
AAC is lossy. The delay/padding behaviour of the encoder means decoded samples won't be bit-identical to the PCM you fed it, and the re-encoded file is a new file with the audio stream changed. Also remember that a trimmed or non-MP4 view can get transcoded by Core's exporter, so don't claim the whole file is untouched after this node.
On audio-only graphs, use the sibling AUDIO node. There's a matching envelope for the AUDIO socket in this pack that takes the frame rate as a string instead of reading it from the video component. Same idea, different wire.
Inputs (4)
| Name | Type | Default | Description |
|---|---|---|---|
| video | VIDEO | — | |
| enabled | BOOLEAN | true | — |
| mute_first_frames | INT | 1 | 对应开头N帧的声音置零,保留时长;太大会削掉首字。 |
| fade_in_ms | FLOAT | 100–10000 | 静音后半余弦淡入;N=0时仅淡入。不做全片降噪或归一化。 |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| video | VIDEO | — |
| report_json | STRING | — |