Nodes/comfyui-minimax-h3-audio-T8/T8 开头音频静音+淡入(AUDIO)
ComfyUI Node

T8 开头音频静音+淡入(AUDIO)

Kill it without touching the timing

By T8mars·Created 2 months ago·Updated 2 days ago· 1,157
T8 开头音频静音+淡入(AUDIO)
  • audio
  • audio
  • report_json
◄enabledtrue►
◄mute_first_frames1►
◄fade_in_ms10►
◄fps24►

Generated H3 audio frequently starts with a transient: a pop, a blip, half a breath, or a syllable that arrives before the picture does. The obvious fix - trim the head of the clip - breaks lip sync on everything downstream. This node is the fix that doesn't move the timeline.

What it does

MiniMaxH3AudioOpeningMuteFadeT8 zeroes the first N frames' worth of samples and then applies a half-cosine fade-in, in place, keeping the total sample count identical. Nothing is shortened, nothing is resampled, no time shift. It's an envelope pass over the waveform and that's the whole job.

That's a narrow tool, and narrow is what you want here. It runs on the AUDIO type directly, so it slots between whatever produced your track and whatever consumes it - an H3 audio output, a VAE decode, an audio-driven avatar chain, a saver.

The two knobs that matter

mute_first_frames (default 1) is frames of video time, not samples: the node converts through fps at your sample rate. That conversion is why the fps field is a string with a default of "24" and why the tooltip nags that it must match the video - fractional rates like 30000/1001 are accepted. Set fps wrong and you mute the wrong duration by a measurable margin, which is the single easiest way to screw this node up.

fade_in_ms (default 10) is the ramp after the mute, half-cosine. The author's own tooltip explains why half-cosine rather than linear: it's continuous where it meets the muted region, so you don't trade a click for a small step.

enabled is a bypass, and it's real: with enabled off you get the input tensor back untouched. Worth knowing because mute_first_frames = 0 plus fade_in_ms = 0 also short-circuits to a no-op, which is the cleanest way to A/B the setting in place.

The outputs are audio and report_json. The report is the useful part nobody reads until something sounds wrong: it records mute_samples, fade_samples, requested_fade_samples, mute_end_seconds and fade_end_seconds as exact fractions, so you can prove what the envelope actually did instead of squinting at a waveform. The pack's sibling diagnostics node will read it if you want it in English prose.

Installing

ComfyUI Manager → MiniMax H3 Audio T8, or:

cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8

Restart ComfyUI fully and refresh the page. Package-wise this node needs nothing: requirements.txt in the pack installs zero extra packages, and the base nodes lean on the torch/numpy/Pillow/safetensors that ComfyUI already ships. This one doesn't even need a model loaded - it's pure tensor math, so you can test it on any AUDIO input while a generation is running.

Where it bites

Too many muted frames eats the first word. The schema says it plainly: 太大会削掉首字 - go too big and you clip the opening syllable. The default of 1 frame (about 42 ms at 24 fps) is chosen to be below speech onset. If you need 4-5 frames to kill a transient, the transient is probably the bigger problem.

It is not a cleanup pass. No denoising, no normalization, no loudness matching. If your H3 audio opens loud and settles, this won't fix the level; it'll just silence the first 40 ms of it.

PCM-level guarantee only. The node's report carries a limitation field stating the envelope is exact at the PCM contract - once the track goes through a lossy encoder, decoded samples outside the envelope can shift. If you're delivering AAC, judge the result from the final file, not from an intermediate waveform view.

One real edge case, handled deliberately. A one-sample fade envelope evaluates to zero, and the code chooses continuity at the mute boundary over having a technically nonzero ramp. So a very short fade_in_ms at a low sample rate collapses to a hard mute. That's intended, but it means 1 ms isn't meaningfully different from 0 ms - don't tune in single milliseconds and expect audible change.

CategoryT8/MiniMax H3/Audio

Inputs (5)

NameTypeDefaultDescription
audioAUDIO—
enabledBOOLEANtrue—
mute_first_framesINT1对应开头N帧的声音置零,保留时长;太大会削掉首字。
fade_in_msFLOAT100–10000静音后半余弦淡入;N=0时仅淡入。不做全片降噪或归一化。
fpsSTRING24必须与视频一致;支持30000/1001。

Outputs (2)

NameTypeDescription
audioAUDIO—
report_jsonSTRING—