VRGDG H3 Apply Latent Continuation Guide
Gluing Two H3 Scenes Together Without Touching a Single Pixel
- positive
- latent
- context_latent
- positive
- latent
This is the node that actually performs the continuity trick. VRGDG H3 Load Latent fetches the tail of your previous scene as a tensor; VRGDG H3 Apply Latent Continuation Guide hands that tensor to MiniMax H3 as a native temporal keyframe, so the model conditions on its own real prior latent instead of an image you decoded and re-encoded.
MiniMax H3 is a 33B omni-modal model that generates picture and sound in one pass, and its conditioning accepts keyframes fed as latents, not just pixels. This node exploits that. The usual "make a long video" pattern is decode → extract last frame → encode that frame → condition, which throws away detail at the decoder and picks up whatever the decode hallucinated. This path pays none of that.
How it works
It takes your positive conditioning and appends a keyframe entry to the minimax_keyframes list the H3 conditioning carries (resolved_frame_index plus the latent), which is done through ComfyUI's own conditioning_set_values helper - the same mechanism the pack's image-reference node uses. It appends rather than replaces, so if you already have keyframes stacked up from elsewhere they survive.
Two details in there matter. First, it checks the target latent's spatial size against the context latent, and if they differ it resamples the context with bicubic interpolation (scaling to cover and centre-cropping when the aspect ratio differs by more than 5%). Handy, but the clean answer is to keep scene resolutions identical. Second, include_audio anchors the predecessor's audio tail too - the tooltip is explicit that you leave it off when the scene's audio already comes from a source track (Audio Drive, an audio reference), because the predecessor's audio would collide with it at the same timeline position.
The console line it prints is your best debugging tool: token count, spatial size, whether audio got attached, how many existing keyframes were already there. When a seam looks wrong, read that before you touch anything.
Inputs and outputs
Required: positive (the conditioning you'd otherwise hand your guider - the pack's own builder patches the guider's conditioning socket directly with this node's output), latent (the target latent, used only to inspect the resolution, so wire the same latent your sampler takes), and context_latent (the output of VRGDG H3 Load Latent).
Optional: frame_idx, default 0 - "where the continuation guide starts". Leave it at 0 for plain continuation. In exact-frame mode the guide has to sit before the pinned last-frame image, so the offset is non-zero; the Builder computes it from its warm-up plan, and hand-wired you'd do that arithmetic yourself.
include_audio is the one boolean, default false. See above; the default is right more often than not.
Outputs: positive (into your guider / conditioning input) and latent (a straight passthrough, so you can keep routing the sampler's latent through the node if that reads more cleanly on the canvas). Nothing about the latent is modified - all the work happens on the conditioning side.
The part people skip, and then complain about
The context frames are added to the front of your render. They are not the scene. Your output now starts with the predecessor's last 22-odd frames reproduced, and those have to come off again. In the AI Video Builder this is the warm-up, and the timing plan trims video and audio together by the same amount, which is why lip sync survives. Wire this node by hand and forget the trim, and your final file starts with a duplicate of the previous scene - or worse, you trim the video only and the audio slides out of sync.
That's what VRGDG H3 Trim Continuation Output is for in a manual graph. In the Builder it isn't even used; the timing does it.
Install
Manager → Install Custom Nodes → search vrgamedev, or:
cd ComfyUI/custom_nodes
git clone https://github.com/vrgamegirl19/comfyui-vrgamedevgirl.git
python -m pip install -r comfyui-vrgamedevgirl/requirements.txt
Restart and hard-refresh the browser. Requirements note: the pack installs voxcpm, llama-cpp-python, demucs and av for its Builder and audio features, so on Windows portable with Python 3.13 run the README's bootstrap (upgrade pip/setuptools/wheel, then Cython scikit-build-core) or llama-cpp-python builds from source.
You also need a reasonably current ComfyUI: H3's keyframe conditioning and comfy.nested_tensor (which carries the video+audio pair) are recent additions. On an older frontend the node loads but the guide goes nowhere useful.
When it fights you
Three things. Scenes at different resolutions: the resize logic saves you from a crash, but resampled context is context you paid a quality tax on. A predecessor you re-rendered: the latent on disk is from the old render, so your continuation is continuing something that no longer exists - re-render this scene. Scene 1: there's no predecessor latent, so this mode starts at scene 2 by definition.
Inputs (5)
| Name | Type | Default | Description |
|---|---|---|---|
| positive | CONDITIONING | — | |
| latent | LATENT | — | |
| context_latent | LATENT | — | |
| frame_idxopt | INT | 00–999999 | Frame index where the continuation guide starts (typically 0) |
| include_audioopt | BOOLEAN | false | Also anchor the predecessor's audio tail. Leave off when the scene audio is already driven by a source track (Audio Drive / audio reference); the predecessor audio would conflict with it at the same timeline position. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| positive | CONDITIONING | — |
| latent | LATENT | — |