✂️ Trim Video To Audio (OreX)
The Clip Is Always a Few Frames Longer Than the Audio. This Cuts It Back.
- video
- audio
- video
- duration_sec
Every video model does the same petty thing: it refuses your frame count and rounds it up to a number it likes. Wan wants 4n+1, so 78 frames becomes 81. You asked for a five-second shot, you got 5.06 seconds, your audio is still 5.02. One clip, nobody notices. Four clips joined into a scene, and you get desync and a silent tail at every seam.
Trim Video To Audio fixes exactly that. It doesn't trim by a timestamp you type in - it takes the length from the audio and rebuilds the video to match.
How it works
It runs on ComfyUI's native VIDEO type, not on files: get_components() gives back decoded frames as an [N, H, W, C] tensor plus frame_rate, the audio as the usual waveform/sample_rate pair. Then:
n_frames = ceil(audio_duration * fps) # with a 1e-3 frame tolerance
That tolerance matters: a float artefact landing on 102.0000001 frames would otherwise round up to 103 and hand you back the silent tail. The result is clamped so you never get zero frames, and a video with fewer frames than the audio needs is not stretched. Then the audio is cut or padded to n_frames / fps samples and both come back as a fresh video.
No ffmpeg, no temp files: the sibling nodes in this pack (Advanced Video Load, Video Preview) shell out to ffmpeg, this one just slices tensors. It also replaces the incoming clip's audio with the track you wire in; it doesn't mix.
The three inputs, and the one that matters
video (VIDEO) and audio (AUDIO) are self-explanatory; the useful setting is audio_fit, which only ever affects the short side:
- pad_silence (default) - audio kept intact, silence appended up to the frame boundary, at most one frame's worth. Nothing in the voice-over is lost, which is what you want when you're assembling clips afterwards.
- trim - the audio is never extended. Longer than the video, it gets cut to the video's length.
One caveat straight from the source, which the README doesn't spell out: the truncation branch runs in both modes. pad_silence protects an audio track a hair shorter than the frame boundary - it does not protect a track genuinely longer than the footage you have. There the tail gets cut either way, because the node refuses to invent frames.
Outputs are video (VIDEO) and duration_sec (FLOAT). Wire the video into a Save Video, this pack's Video Preview, or core's Concatenate Video. duration_sec is the final length rounded to the frame boundary: a sanity readout, handy for stamping into a filename. This is not an output node, so keep your save node connected.
Installing it
ComfyUI Manager, search comfyui-OreX (the pack ships as "Orex Nodes"). Or:
cd ComfyUI/custom_nodes
git clone https://github.com/orex2121/comfyui-OreX
# restart ComfyUI
The pack's requirements.txt pulls in Pillow, soundfile, pydub, pyloudnorm and modelscope - modelscope is for the Skin Retouching node, not this one, so you're installing a 27-node pack with a heavy dependency list to get a ninety-line node. The node itself imports only math, torch and comfy_api. What it does need is a reasonably current ComfyUI: it imports from comfy_api.latest, with a fallback to the older comfy_api paths, so a truly ancient install fails at load with an ImportError. Update first, then blame the node.
Where people get burned
The whole calculation hangs on the frame_rate the VIDEO object carries, and there's no widget to correct it. If something upstream rebuilt your frames from an IMAGE batch with the wrong fps - classic after a RIFE pass where you interpolate 16→32fps and forget the Create Video widget - the node trims confidently to the wrong boundary. The math is right; the fps is lying.
Second: audio needs a real AUDIO - core Load Audio, this pack's Audio Load, or GetVideoComponents on the same clip, whose audio output lets you round a video down to the length of its own embedded track.
And if you're joining clips inside ComfyUI, core's Concatenate Video has a complete_audio input that overrides whatever audio the segments carry - in that path the desync never reaches the file anyway. This node earns its place when each clip has to stand on its own: a Save Video deliverable, or files handed to DaVinci or an ffmpeg concat where every segment needs its own correctly-sized track.
The verdict: a plain, single-purpose node doing one annoying thing properly, costing nothing to run. Reach for it as the last step of a TTS → video → assembly pipeline; if you're just cutting seconds off the front of a clip, core's Video Slice already does that. The pack is obscure - the OreX handle has essentially no reddit footprint - and the author's README opens with "delete the nodes from custom_nodes and install them again." Take that as the maintenance advice it is if an update leaves you with a broken import.
Inputs (3)
| Name | Type | Default | Description |
|---|---|---|---|
| video | VIDEO | Видео для обрезки / Video to trim | |
| audio | AUDIO | Аудио, по длине которого обрезаем видео / Audio whose length defines the result | |
| audio_fit | COMBO | pad_silence | pad_silence: аудио не режется, в конец добавляется тишина до границы кадра. trim: аудио режется по длине видео, тишина не добавляется / pad_silence: audio is kept, silence is added up to the frame boundary. trim: audio is cut to the video length, no padding |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| video | VIDEO | — |
| duration_sec | FLOAT | — |