Video Continuation Concat
How you hang the last 22 frames on the next clip
- prefix_images
- prefix_mask
- prefix_audio
- body_images
- body_mask
- body_audio
- images
- mask
- audio
- trim_info
This is the node in the middle of the chain. Something upstream hands you the tail of the previous segment; something downstream is generating the next one; this node composes the two into a single timeline, produces the mask that says which part is context and which part is new, and - this is the part people miss - emits a small metadata token that the rest of the chain uses to undo the composition afterwards.
That token is why the pack works the way it does. Chaining by hand means carrying "we prepended 22 frames, remember to cut them off" in your head across four nodes. Here it's a value on a wire.
concat vs replace
mode has exactly two options and picking the right one is the whole configuration.
concat (the default) prepends prefix_images to body_images. Your timeline gets longer: 22 frames of context plus however many you're generating. The sampler sees the prefix as extra frames, which is the simple, honest version of continuation and the one to start with.
replace keeps the body's length and overwrites its leading slots with the prefix - nothing grows, because the 22 context frames were already in the body timeline. Use it when your target latent already contains those slots, which is the shape H3's continuation layout tends to produce. A prefix longer than the body is an error; there'd be nowhere to put it.
Either way the prefix is temporary. Both modes mark it for trimming, and trim_info carries trim_frames so Trim Video Continuation Prefix removes exactly the same frames the sampler generated over.
The mask contract
Two defaults, and they're chosen so that an unconfigured node does the sane thing:
- No
prefix_mask→ the prefix gets an all-zero mask. Zero means preserve, so the context frames are held, not redrawn. This is normally what you want: the prefix is reference material, not paint. - No
body_mask→ the body gets an all-one mask. One means redraw, so the body is what the sampler is asked to work on.
Both tooltips are explicit that these are video/image masks and not audio masks - audio preservation is a separate job, handled downstream by H3 Set Audio Prefix Noise Mask. The mask output is a per-frame redraw mask ready for the pack's Set Video Latent Noise Mask node (with type set to minimax for H3), which converts it into the latent-shaped mask a sampler consumes.
Masks accept [H,W] or [frames,H,W]; a single-frame mask is expanded across time, and a wrong-resolution mask is nearest-resized to the frames.
Inputs and outputs
frame_rate (default 24) and mode are the two required widgets. body_images is nominally optional but required in practice - the node raises body_images is required if you leave it empty. prefix_images comes from Load Indexed Video Segment or any image batch; it must match the body's height and width, and there's an explicit error if it doesn't.
prefix_audio and body_audio are optional and their absence is meaningful rather than an error: a missing waveform is filled with duration-matched silence. Supply prefix_audio when you have real audio to carry forward; leave body_audio empty when H3 generates the body's sound. prefix_mask or prefix_audio without prefix_images is an error - there's nothing to apply them to.
Outputs are images (the composed batch), mask (per-frame redraw), audio (the composed waveform, prefix and body resampled to one sample rate, with the prefix boundary recorded as an exact sample count), and trim_info. Send images and mask toward your latent and sampler, audio wherever it's going, and split trim_info to both Trim Video Continuation Prefix and H3 Set Audio Prefix Noise Mask.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/wjie98/comfyui-svdint4.git
# restart ComfyUI
Same pack under two names: GitHub and Manager say comfyui-svdint4, the README says "ComfyUI Turing Utils" and clones comfyui-turing-utils. No models to download. The CUDA kernel build the README walks you through is only for the ConvRot loaders and attention nodes - this one is pure Python plus PyAV, which ComfyUI already has.
Worth knowing before you go looking for help: I searched the community corpus for this pack, its author, and both repo names and found nothing. There's no thread with your answer in it yet.
Where it bites
Mode confusion is the classic. Concat then trim is a growing timeline; replace then trim is the same length you started with. Trim both, and if your output is 22 frames longer than you expected, you ran concat and forgot the trim.
Nothing here is audio-aware. It composes waveforms, but which part of the audio the sampler should regenerate is decided by the audio latent noise mask. Wire the trim metadata to that node or H3 will happily write a new soundtrack over your carried-over audio.
And the prefix is only as good as what you feed it. A clean 22-frame tail gives real continuity; a prefix that's already been through Video Prefix Context Noise tells the model "this is history, don't copy it". Both are legitimate - they just produce visibly different seams.
Inputs (8)
| Name | Type | Default | Description |
|---|---|---|---|
| frame_rate | FLOAT | 24.000.01–1000 | — |
| mode | COMBO | concat | 2 options: concat, replace |
| prefix_imagesopt | IMAGE | — | |
| prefix_maskopt | MASK | Video/image redraw mask for the prefix; this is not an audio mask. | |
| prefix_audioopt | AUDIO | Optional prefix waveform content. Audio preservation is controlled later by the H3 audio latent noise mask. | |
| body_imagesopt | IMAGE | Required generated or source body frames. | |
| body_maskopt | MASK | Video/image redraw mask for the body; this is not an audio mask. | |
| body_audioopt | AUDIO | Optional body waveform content. Leave empty when H3 should generate the body audio. |
Outputs (4)
| Name | Type | Description |
|---|---|---|
| images | IMAGE | — |
| mask | MASK | — |
| audio | AUDIO | — |
| trim_info | TURING_UTILS_VIDEO_TRIM_INFO | — |