Nodes/comfyui-svdint4/Merge Indexed Video Segments
ComfyUI Node

Merge Indexed Video Segments

40 MP4s in, one out, no re-encode

By wjie98·Created 3 months ago·Updated 3 days ago· 3
Merge Indexed Video Segments
    • filename
    • segment_count
    ◄root_directoryvideo/segments►
    ◄max_index-1►
    ◄output_filenamemerged.mp4►
    ◄overwritefalse►
    ◄audio_bitrate_kbps192►

    You've got a folder of numbered segments and you want a video. The naive answer - decode all of them into one IMAGE batch and run them through a Video Combine - is a bad trade in every dimension: it re-encodes everything, it holds the whole thing in memory, and it hands the audio over to whatever the encoder does at a seam. Twenty junctions later, your sound drifts against the picture.

    This node does the other thing. It stream-copies the compressed video and re-encodes the audio exactly once.

    How it works

    The video path never touches pixels. Each segment's packets are demuxed and muxed into the output with their pts/dts shifted by the segment's exact starting frame, converted through the file's own time base - and it refuses to proceed if that time base can't represent the boundary exactly rather than rounding and letting the error accumulate. Timestamps are also checked to be strictly increasing as it goes, so a segment with broken timing fails loudly at the point of the problem instead of producing a subtly stuttery output.

    The audio path is where the real work is. Every AAC segment carries its own independent encoder delay and padding, so "just concatenate the audio" is wrong by a few milliseconds per join - and a few milliseconds per join is how a 20-segment video ends up a quarter-second out of sync. Instead, each segment's audio is decoded, its padding removed, the sample count fitted to that segment's exact cumulative frame boundary (start_frame * rate / frame_rate, in rationals), and the whole track is encoded once at the end as a single continuous stream. Legacy segments that decode up to one AAC frame short get padded automatically; anything larger is an error, which is the right call.

    Inputs and outputs

    root_directory is the folder of NNNNNN.mp4 files, same resolution rules as the rest of the pack (relative → under ComfyUI's output directory, absolute → used directly). max_index is inclusive and defaults to -1, meaning every segment through the highest existing index - that's the one you want for a finished chain. 0 merges only 000000.mp4, which is handy for testing.

    output_filename defaults to merged.mp4 and must be a bare filename ending in .mp4. It explicitly refuses the reserved NNNNNN.mp4 form, so you can't accidentally create a file the loader and merger will treat as a segment. overwrite defaults to false - same guardrail as the saver, since re-running a merge is a two-second operation but blowing away your only good output is not. audio_bitrate_kbps (advanced, default 192) is the one knob on the audio re-encode. If your chain carried dialogue, 192 is fine; if it's music or anything dense, 320 costs you nothing at these lengths.

    Outputs are filename and segment_count, and the merger is an output node with no preview attached. segment_count is the useful one: it's your receipt that twenty-eight files actually went in.

    Install

    cd ComfyUI/custom_nodes
    git clone https://github.com/wjie98/comfyui-svdint4.git
    # restart ComfyUI
    

    It's listed as comfyui-svdint4, its README titles the project "ComfyUI Turing Utils" and clones comfyui-turing-utils - renamed repo, same pack. You do not need the pack's CUDA kernel build for this node; it's PyAV and torch, and PyAV ships with ComfyUI. The pack's requirements.txt adds only safetensors.

    I searched the community corpus for the pack and its author and found nothing at all, so there's no thread where someone has already hit your problem. The README is dense and accurate, which makes up for most of it.

    Where it bites

    Contiguity is mandatory. It checks 000000.mp4 upward and names the missing files if there's a gap. Delete a bad segment and you must renumber the rest, not leave a hole.

    Compatibility is checked, hard. All segments must share a frame rate and a video signature - codec, dimensions, pixel format, colour metadata. This is why a merge of segments from two different graphs or two different CRF settings can fail: they're not the same video stream. The error message names the offending file.

    Audio presence must match across the folder. Some segments with audio and some without is an error, not a guess.

    And it copies what's stored. That's the headline promise and the biggest trap in the same sentence: if the segments still contain their 22-frame continuation prefixes, those frames are now part of your film, repeated at every junction. Trim Video Continuation Prefix goes before the saver, every time. If you've already saved untrimmed segments, the fix is to re-trim and re-save them, not to hope the merger is clever.

    CategoryTuring Utils/video

    Inputs (5)

    NameTypeDefaultDescription
    root_directorySTRINGvideo/segments—
    max_indexINT-1-1–999999Last segment index to include, inclusive. -1 merges every segment; 0 merges only 000000.mp4.
    output_filenameSTRINGmerged.mp4—
    overwriteBOOLEANfalse—
    audio_bitrate_kbpsINT19232–512—

    Outputs (2)

    NameTypeDescription
    filenameSTRING—
    segment_countINT—