H3 Extend Append
Joining the segment back onto the clip without a seam
- latent
- sampled
- latent
- frames
- report
Chaining video segments by hand has one job that always gets done badly: the join. You generate a continuation, and it comes back with the frames you carried inside it as well as the new ones. Cut in the wrong place and you either duplicate a beat or drop one, and the clip stutters. Do it in decoded pixels and you've paid a VAE round trip to find out.
H3 Extend Append is the join, done in latent space, in one node. It's the second half of the pair: H3 Extend Window opens a segment, your sampler fills it, this node drops the overlap off the front and appends what's left to the clip so far.
What it does with the overlap
latent is the clip so far - the loop's carried value. sampled is whatever the sampler returned for the window H3 Extend Window opened. overlap_frames is the number that window reported, and it's the only thing that needs to match between the two nodes; wire the output to this input and you can't get it wrong.
On the first pass (segment_index 0, and it's optional) there's nothing to join to, so what came back is the clip. Which means you can run the same two-node loop for segment one and never special-case it.
After that, the sampled segment arrives with the clip's own last frames at its head. The node drops exactly that many frames and appends the rest - every earlier frame is carried through untouched, because nothing re-encoded it.
A cut is the interesting case, and the one that shows the node knows about H3's grid. With an overlap of 0, the whole new scene is appended after the clip's last full 17-frame block, trimming the 5 frames past it. H3's frame counts sit on a 17k+5 grid, so this is what keeps a cut landing on a legal length instead of half a block off.
Outputs: latent - the longer clip, for the next pass or for a decode at the end. frames - how many frames the joined clip now holds, which is the number you want on screen. And report - what was carried and what was added, in words, so a length that looks wrong is diagnosable without guessing.
Driving it
The intended shape is a While Loop. Its index goes to segment_index here and to pass_index on H3 Extend Window; the loop's carried value is the clip, which enters latent, exits latent, and gets stored back on each iteration. Prompt rows come from MiniMax H3 Conditioning, either through the window node's prompts input or as plain conditionings.
Decode once, after the last segment. This is the instruction people ignore, and they pay for it: decode every pass and you're VAE-decoding the whole clip over and over, each decode a lossy round trip, and the joins show up as small colour shifts between segments. Decode the finished latent and none of that happens, because the join was never in pixels.
Installing it
Ships in WAS Node Suite v3 - MIT, by WASasquatch, 468 nodes across images, masking, text, latents, sampling, files, animation and video. Install the pack, not the node:
cd ComfyUI/custom_nodes
git clone https://github.com/WASasquatch/was-node-suite-comfyui.git
ComfyUI Manager, search WAS Node Suite v3, is the route the README recommends. Needs ComfyUI 0.14.0+ and Python 3.10+. The pack installs no Python packages, downloads nothing, and never runs pip; its config.yaml and folders appear under <ComfyUI user dir>/was-node-suite/ on first start, which is why that first boot is a second or two slower and an update repeats it. Anything with "was-node-suite-comfyui" and Cannot import module in the error is a v2-era problem - twenty pip packages the v3 rewrite deliberately dropped - and reinstalling will not fix a copy that old.
When it complains
Three failures, all specific and all from the node's own error paths.
"Not an H3 joint latent" - you're feeding it something that isn't H3 video-plus-audio. An empty latent from a plain Empty Latent Image won't do; H3 wants its joint 24-channel video and audio latent, which is what the conditioning node and the window produce.
The clip is shorter than the overlap - the carried value went in empty on a pass that isn't the first, usually because the loop's initial value wasn't seeded with segment one's latent.
The pass added no frames - H3 Extend Window's overlap ate the entire sampled segment. Raise extension_frames relative to the overlap; a 22-frame carry over a 17-frame extension leaves nothing to append.
Inputs (4)
| Name | Type | Default | Description |
|---|---|---|---|
| latent | LATENT | The clip so far, from the loop's carried value. | |
| sampled | LATENT | What the sampler returned for the segment H3 Extend Window opened. | |
| overlap_frames | INT | 220–362 | The overlap H3 Extend Window reported, `0` for a cut. |
| segment_indexopt | INT | 00–64 | Which segment this is, from `0`. Wire the same While Loop Open index that drives H3 Extend Window. |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| latent | LATENT | The longer clip, for the next pass or for a decode. |
| frames | INT | Frames the joined clip now holds. |
| report | STRING | What was carried and what was added. |