H3 Continuity • Append & Stage
Extend the take instead of re-rendering it
- sampled
- context
- cumulative_latent
- ticket
MiniMax H3 shoots 4–15 seconds at 24 fps with audio welded to the picture, and the moment you get a take you like, you want thirty seconds of it. The usual answer is to generate three clips and stitch the mp4s, which is exactly how you get a cut, a costume change and a room-tone jump at seconds 14 and 29. Append & Stage is the other answer: it joins the latents, not the video files.
It's a two-wire utility between your sampler and your decoder that welds the pinned previous clip onto the new sample, then writes the result to disk as a checkpoint. Deliberately, it isn't usable on its own: both inputs come from the pack's MiniMax H3 Director pipeline.
What it actually does
The interesting thing here is where the seam lives. ComfyUI's chunking helpers reach 20–30 seconds by hiding the stitch, and identity drift across chunk boundaries survives it unchanged - the same wall Wan people hit at 81 frames. H3 continuation instead conditions the new segment on the tail of the previous latent, samples a fresh window, and then concatenates at the token level, so the source prefix is exact at the latent level rather than re-encoded.
Concretely, the context input carries the session, a run ID, the mode, the resolved prompt and a small token layout. If continuity is active, the node loads the pinned parent out of output/df_h3_continuity/<session>/<clip_id>/latent.safetensors, canonicalises dtype/device on both the video and audio streams, and appends the new sample onto it. If continuity is off, it hands your sample straight through untouched.
Short version: the sampler renders an overlapping context window at the head of the sample - 5, 22, 39, 56 or 73 frames - which the prompt asks the model to "weld" to. Those frames are discarded before the append. They exist so the model has real motion, camera and sound to match against, and they're why visible duration grows only by the new, multiple-of-17 extension, not by the whole sampled window.
Then it stages. The cumulative latent goes to latent.safetensors plus a clip.json marked staged, with frame count, 24 fps, canvas, model family, parent ID and prompt - and you get back a ticket. A staged checkpoint is not a source. Nothing will offer it to you as a "latest output" until the export proves itself, which is the other node's job.
Inputs and outputs
sampled(LATENT) - the sampler's raw output, i.e.SamplerCustomAdvanced.output. Not an image, not a decoded video.context(DF_H3_CONTINUITY_CONTEXT) - from the Director Guide'scontinuity_contextoutput. This is the only place that type exists.cumulative_latent(LATENT) - source prefix plus new audio/video tokens; wire it into the latent upscale and AV decode you already have.ticket(DF_H3_CONTINUITY_TICKET) - goes to Publish Export, and nowhere else. Don't try to bridge it.
There are no optional inputs and no widgets. That's the whole node.
Director Guide ──context──► Append & Stage ──cumulative_latent──► upscale ──► AV decode
SamplerCustomAdvanced ──sampled──┘ └──ticket──► Publish Export.ticket
Video exporter ──filename──────────────────────────────────► Publish Export.filename
Install
The pack is one install for all of it: ComfyUI Manager → search DaSiWa-Nodes, or manually:
cd ComfyUI/custom_nodes
git clone https://github.com/darksidewalker/ComfyUI-DaSiWa-Nodes
pip install -r requirements.txt
# restart ComfyUI
Two things in that requirements file matter here. Continuity uses PyAV (av>=18.0) to probe and import local media - no external ffmpeg/ffprobe binary needed. And the Tail Guide path requires a ComfyUI build with native arbitrary-frame H3 guides; an older checkout raises "requires ComfyUI commit e01fb4c or newer", so update ComfyUI before filing a bug. The pack also pulls in nvidia-vfx and llama-cpp-python for nodes you're not using here, but __init__.py imports every module up front and custom nodes share one Python environment with no isolation - install the requirements file rather than cherry-picking.
Also worth saying out loud: H3's weights are under a community licence whose applicable territory excludes the US, EU, UK and South Korea. The pack doesn't enforce anything; the licence is on you.
Where it bites
The classic "this node is broken" moment: it passes your sample straight through with a disabled ticket, and Publish Export prints "Continuity capture is off." You never selected a start video or a completed checkpoint. Capture being enabled isn't the same as continuation being active - picking a source is what does it.
Second, everything must line up. The source has to be native 24 fps, the same model family (REF2VA and first/last-frame modes don't mix) and the same canvas dimensions, or the Director refuses before sampling. Choose a new visible duration via the Director's Duration field rather than hand-editing frames - the valid range is up to 15 seconds, and the combined context + new window has to stay inside the native 362-frame limit.
Finally, don't treat the export as the clip - the latent is authoritative, so a downstream re-decode can change how the already-finished part looks, and ping-pong or trimming on the export breaks validation outright. Save the workflow too: the selected source, next-action prompt and Duration live in the graph's timeline data, not in this node.
Inputs (2)
| Name | Type | Default | Description |
|---|---|---|---|
| sampled | LATENT | — | |
| context | DF_H3_CONTINUITY_CONTEXT | — |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| cumulative_latent | LATENT | — |
| ticket | DF_H3_CONTINUITY_TICKET | — |