Nodes/comfyui-svdint4/Load Indexed Video Segment
ComfyUI Node

Load Indexed Video Segment

The number is the whole trick

By wjie98·Created 3 months ago·Updated 3 days ago· 3
Load Indexed Video Segment
    • images
    • audio
    • frame_rate
    ◄root_directoryvideo/segments►
    ◄segment_index0►
    ◄tail_frames22►

    Every long AI video you've seen that isn't one shot is a chain of shorter generations glued together, and the glue is where the quality dies. The Wan-era version of this was "generate clip one, feed its last frame into clip two" - and the community verdict on it is blunt: identity drift across chunk boundaries, forever.

    This pack takes a different tack. Instead of a loop inside one node, it keeps every segment on disk as 000000.mp4, 000001.mp4, 000002.mp4 … and gives you a loader that knows how to pick up where the last one left off. The numbering isn't cosmetic. It's the state machine.

    The index convention, which is the actual feature

    segment_index is a required input with a default of 0, and the three cases behave differently on purpose:

    • 0 returns empty outputs. Not the first segment - nothing. That's the first pass of your chain, which has no previous clip to continue from.
    • i loads segment i-1. So if the saver downstream is writing segment 3, this loader is reading 000002.mp4. One number drives both halves of the loop, which is what makes a chained workflow runnable repeatedly without you renumbering anything. (A literal i semantics here would be off-by-one hell.)
    • -1 loads the highest six-digit MP4 in the directory. This is the resume button. Crash on segment 27, set -1, and you're continuing from what's actually on disk rather than what you meant to render.

    root_directory defaults to video/segments and is resolved below ComfyUI's own output directory when relative - a relative path that escapes upward is rejected. Absolute paths are used as-is, which is what you want if your segments live on a scratch drive.

    tail_frames defaults to 22 and that number is not arbitrary: 22 is 17+5, the continuation prefix H3 wants, in its 17*n+5 frame grid. The tooltip says the rest - 0 loads the complete segment, any positive value decodes only the final N frames. You almost never need the whole segment; you need the handle to hold on to.

    Outputs

    images is the tail as an IMAGE batch, ready for Video Continuation Concat's prefix_images socket. audio is the matching AUDIO, already sliced to match those frames: the loader works in exact rationals from the container's real frame rate, so the audio boundary lands on a sample boundary rather than "close enough". frame_rate is a FLOAT taken from the file itself - wire it into the saver and the concat node rather than typing 24 by hand. If any segment in your folder was written at a different rate, the merger will refuse the set, and a float you retyped is the usual culprit.

    A missing file returns empty outputs (None, None, 0.0) rather than an exception. Combined with segment_index 0, that's what lets a fresh run start clean and produce its own first segment.

    Install

    cd ComfyUI/custom_nodes
    git clone https://github.com/wjie98/comfyui-svdint4.git
    # restart ComfyUI
    

    The repo is listed as comfyui-svdint4 but its README titles the project "ComfyUI Turing Utils" and clones comfyui-turing-utils - same pack, renamed along the way. Nothing to download, no model files, and no CUDA build needed: the pack's expensive bit (python -m pip install -v --no-build-isolation -e ./kernel) is for its quantisation and attention nodes, while this one only needs PyAV and torchaudio, both of which ComfyUI already installs.

    I checked the reddit corpus for this pack, the author, and both repo names. Nothing - zero hits. You're on your own with the README, which is unusually detailed, so read it.

    Where it bites

    The file-name check is strict: exactly six digits, ASCII, .mp4, ten characters, a real file. 23.mp4, 000023.MP4, and symlinks are all invisible to it. If the loader says a segment doesn't exist when you can see it in Explorer, look at the name before you look at the path.

    Because it only decodes the tail, tail_frames decides what H3 gets to see of the previous clip. Bumping it (say 30 or 39) buys a little more continuity and costs decode time and context; dropping it to a handful gives the model almost nothing to anchor to. Start at the default 22 and move it deliberately.

    And remember the loop's spine: this node reads, Video Continuation Concat hangs the frames on the front, the sampler generates, Trim Video Continuation Prefix removes the temporary context, Save Indexed Video Segment writes the next number. If you skip the trim, you'll save a segment that already contains the previous one's tail - and Merge Indexed Video Segments copies what's stored, so those duplicated frames land in your final file.

    CategoryTuring Utils/video

    Inputs (3)

    NameTypeDefaultDescription
    root_directorySTRINGvideo/segments—
    segment_indexINT0-1–1000000—
    tail_framesINT220–16384Final frames to load; 0 loads the complete segment.

    Outputs (3)

    NameTypeDescription
    imagesIMAGE—
    audioAUDIO—
    frame_rateFLOAT—