Load Indexed Video Segment
The number is the whole trick
- images
- audio
- frame_rate
Every long AI video you've seen that isn't one shot is a chain of shorter generations glued together, and the glue is where the quality dies. The Wan-era version of this was "generate clip one, feed its last frame into clip two" - and the community verdict on it is blunt: identity drift across chunk boundaries, forever.
This pack takes a different tack. Instead of a loop inside one node, it keeps every segment on disk as 000000.mp4, 000001.mp4, 000002.mp4 … and gives you a loader that knows how to pick up where the last one left off. The numbering isn't cosmetic. It's the state machine.
The index convention, which is the actual feature
segment_index is a required input with a default of 0, and the three cases behave differently on purpose:
0returns empty outputs. Not the first segment - nothing. That's the first pass of your chain, which has no previous clip to continue from.iloads segmenti-1. So if the saver downstream is writing segment3, this loader is reading000002.mp4. One number drives both halves of the loop, which is what makes a chained workflow runnable repeatedly without you renumbering anything. (A literalisemantics here would be off-by-one hell.)-1loads the highest six-digit MP4 in the directory. This is the resume button. Crash on segment 27, set-1, and you're continuing from what's actually on disk rather than what you meant to render.
root_directory defaults to video/segments and is resolved below ComfyUI's own output directory when relative - a relative path that escapes upward is rejected. Absolute paths are used as-is, which is what you want if your segments live on a scratch drive.
tail_frames defaults to 22 and that number is not arbitrary: 22 is 17+5, the continuation prefix H3 wants, in its 17*n+5 frame grid. The tooltip says the rest - 0 loads the complete segment, any positive value decodes only the final N frames. You almost never need the whole segment; you need the handle to hold on to.
Outputs
images is the tail as an IMAGE batch, ready for Video Continuation Concat's prefix_images socket. audio is the matching AUDIO, already sliced to match those frames: the loader works in exact rationals from the container's real frame rate, so the audio boundary lands on a sample boundary rather than "close enough". frame_rate is a FLOAT taken from the file itself - wire it into the saver and the concat node rather than typing 24 by hand. If any segment in your folder was written at a different rate, the merger will refuse the set, and a float you retyped is the usual culprit.
A missing file returns empty outputs (None, None, 0.0) rather than an exception. Combined with segment_index 0, that's what lets a fresh run start clean and produce its own first segment.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/wjie98/comfyui-svdint4.git
# restart ComfyUI
The repo is listed as comfyui-svdint4 but its README titles the project "ComfyUI Turing Utils" and clones comfyui-turing-utils - same pack, renamed along the way. Nothing to download, no model files, and no CUDA build needed: the pack's expensive bit (python -m pip install -v --no-build-isolation -e ./kernel) is for its quantisation and attention nodes, while this one only needs PyAV and torchaudio, both of which ComfyUI already installs.
I checked the reddit corpus for this pack, the author, and both repo names. Nothing - zero hits. You're on your own with the README, which is unusually detailed, so read it.
Where it bites
The file-name check is strict: exactly six digits, ASCII, .mp4, ten characters, a real file. 23.mp4, 000023.MP4, and symlinks are all invisible to it. If the loader says a segment doesn't exist when you can see it in Explorer, look at the name before you look at the path.
Because it only decodes the tail, tail_frames decides what H3 gets to see of the previous clip. Bumping it (say 30 or 39) buys a little more continuity and costs decode time and context; dropping it to a handful gives the model almost nothing to anchor to. Start at the default 22 and move it deliberately.
And remember the loop's spine: this node reads, Video Continuation Concat hangs the frames on the front, the sampler generates, Trim Video Continuation Prefix removes the temporary context, Save Indexed Video Segment writes the next number. If you skip the trim, you'll save a segment that already contains the previous one's tail - and Merge Indexed Video Segments copies what's stored, so those duplicated frames land in your final file.
Inputs (3)
| Name | Type | Default | Description |
|---|---|---|---|
| root_directory | STRING | video/segments | — |
| segment_index | INT | 0-1–1000000 | — |
| tail_frames | INT | 220–16384 | Final frames to load; 0 loads the complete segment. |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| images | IMAGE | — |
| audio | AUDIO | — |
| frame_rate | FLOAT | — |