Nodes/AuK · T8star-Aix/AuK 音频裁剪 / 时长
ComfyUI Node

AuK 音频裁剪 / 时长

The boring Cut node that saves you from the 30-second wall

By T8mars·Created 22 days ago·Updated 4 days ago· 12
AuK 音频裁剪 / 时长
  • audio
  • 裁剪音频
  • 裁剪时长(秒)
  • 裁剪说明
◄start_seconds0.00►
◄end_seconds0.00►

Why a trim node is even in this pack

AuK is Tencent's speech editing and generation model, and the ComfyUI wrapper for it (AuK · T8star-Aix) is a stack of quiet, annoying rules. The loudest one: source/reference audio and generated output are each capped at 30 seconds, independently. Not added together - a 30-second input can produce a 30-second output. Cross that line and the run fails before inference ever starts.

The second rule is subtler and it's the reason this node exists as a standalone thing. AuK's editing tasks validate against the audio it actually receives, not against whatever number is sitting in your duration widget. Before version 2.0.7 that check ran against the stale pre-trim duration, so a 48-second file you'd cropped down to 4 seconds could still get refused. The fix was to make cropping a real, visible step in the graph, with an exact sample count the downstream node can trust.

So no, it isn't exciting. It's the node the author wired into every single one of the 17 example workflows so the 30-second rule stops being your problem.

What it actually does

It slices the waveform tensor by sample index and hands back the same audio, just shorter. No re-encode, no resample, no downmix. The math is:

first = round(start_seconds * sample_rate)
last  = total if end_seconds == 0 else min(total, round(min(end_seconds, duration) * sample_rate))
cropped = waveform[..., first:last]

That min(...) matters - asking for 60 seconds of a 12-second file gives you the whole 12 seconds rather than an error, and end_seconds = 0 means "to the end of the source." It also deliberately skips the pack's mono-normalization path, so channels survive intact: stereo in, stereo out, and PreviewAudio downstream won't flatten it. Sample rate is untouched. And because the cut is computed in samples rather than seconds, the duration it reports is exact - request 0–4s of a 44.1 kHz clip and you get 176400 samples, and the node says so.

Inputs and outputs

Three inputs, all of them required:

  • audio - a standard ComfyUI AUDIO, i.e. the wire out of Load Audio. Wire the untrimmed load here.
  • start_seconds - float, default 0. Where the cut begins.
  • end_seconds - float, default 0, and 0 means "end of file." Leave both at 0 and the node is a pass-through, which is exactly how the shipped workflows ship it.

Three outputs:

  • 裁剪音频 (AUDIO) - the trimmed clip. In every example workflow this goes straight into AuK Generate / Edit's input_audio. You can also fan it out to a PreviewAudio/SaveAudio pair to hear exactly what the model will hear.
  • 裁剪时长(秒) (FLOAT) - the true length in seconds. Handy to log, or to feed a value node.
  • 裁剪说明 (STRING) - a one-line report like 原音频 12.340s → 截取 0.000–4.000s → 实际输入 4.000s;44100 Hz,2 声道. Text is Chinese; the numbers are the useful part.

One Load Audio can feed several trim nodes at once for different segments of the same file - that's the intended pattern for anything longer than half a minute.

Installing it

The node ships with the pack, so install the pack. In ComfyUI Manager search AuK · T8star-Aix and confirm you're getting 2.0.7 or later - the trim node didn't exist before that, and Manager has been known to still offer an older registry build. If it does:

cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/Comfyui-Auk-T8
cd Comfyui-Auk-T8
python -m pip install -r requirements.txt

Use the same Python that runs ComfyUI. Those requirements deliberately do not touch torch or torchaudio.

The trim node needs no checkpoints. The pack's generator does: AuK Base or AuK-Flash plus Qwen2.5-Omni-3B under ComfyUI/models/auk/, or python download_models.py --variant base|flash|all from the node folder. Flash + Qwen is ~18.7 GB; both variants ~25.5 GB. The pack needs ComfyUI ≥ 0.23.0 - it's written against the V3 io.Schema API, so on an older frontend the node simply won't appear.

Where people get tripped up

The error messages are in Chinese. 裁剪开始超出原音频 12.340s means your start point is past the end of the clip - not a broken install. Same for 裁剪范围无效: end must be greater than start.

Trimming doesn't happen for you. Setting end_seconds to 0 on a three-minute podcast sends three minutes to AuK and fails. Crop explicitly first, then hand it over.

Watch the output side too. Slowing speech or pushing emotion to "sad" (1.22× on the official factors) can push a 28-second input past the 30-second output ceiling. Trim the source further, not the result.

Trust the FLOAT output, not your widget. Your 0.01 steps land wherever you clicked; the cut snaps to the nearest sample. The 裁剪时长(秒) value is the one AuK validates against.

One last thing before you build anything commercial on this: the node code is MIT, but the model weights keep their upstream licenses - and Tencent's releases habitually ship under their own community license, which excludes the EU, UK, and South Korea from the grant. Check the model repo, not the node repo.

CategoryAuK · T8star-Aix

Inputs (3)

NameTypeDefaultDescription
audioAUDIO—
start_secondsFLOAT0.00—
end_secondsFLOAT0.00—

Outputs (3)

NameTypeDescription
裁剪音频AUDIO—
裁剪时长(秒)FLOAT—
裁剪说明STRING—