VRGDG H3 Load Latent
Pull the Last 22 Frames of Your Previous Scene Straight Out of Its Latent
- context_latent
- sliced_tokens
Chaining MiniMax H3 scenes in pixel space is a loop of decode, screenshot, re-encode, and hope. VRGDG H3 Load Latent skips that loop: it opens the predecessor scene's saved latent file and hands you the tail of it as temporal context, exactly as the sampler produced it.
What it actually does
You point it at a project folder and a scene number. It reads <project_folder>/latents/scene_NNN.latent (safetensors, with a torch fallback), pulls out the video tensor and the audio tensor if one was stored, and slices the video along the time axis to keep only the trailing context window. It then repacks video and audio into a ComfyUI nested tensor so H3's joint audio-video sampling still sees a complete latent, and hands it back. No VAE anywhere in sight.
The audio slice is a detail worth knowing, because it isn't just "the same number of frames." The audio latent runs at roughly 40 steps per second against 24 fps video, and the node scales the window to match. So the context isn't only visual continuity - it's the room tone and the tail of whatever was being said.
The inputs that matter
project_folder and scene_number are obvious: the predecessor scene is normally scene_number - 1, and the node defaults to 1, which is why your first scene can't use any of this. Empty folder raises a ValueError; a folder without the file raises a FileNotFoundError that prints the exact path it looked for. The message says "Render the predecessor scene first," and yes, that is the fix.
context_frames defaults to 22 and the tooltip offers the four sensible values: 16, 22, 39, 56. Bigger context = a stronger grip on motion and continuity, and a longer warm-up you have to trim off afterwards. Start at 22.
exact_frame_mode is the fiddly one, and off by default correctly. H3 builds video in token-sized bites (about 1 frame on the first token, then 4 each) and a render can carry extra frames past the scene's real end for alignment or cool-down, so the file's "last frame" isn't necessarily the one you want to continue from. Turn this on and the node cuts the context so it stops just before the predecessor's real last visible frame, deliberately leaving that frame out - you supply it as an image instead, via VRGDG H3 Load Exact Last Frame. It needs the predecessor's latent to have recorded its tail padding; save it with padding unknown and the node silently falls back to plain slicing.
Outputs
context_latent goes into VRGDG H3 Apply Latent Continuation Guide. That's the wire that matters.
sliced_tokens is the sanity check - an integer, not a tensor. A 22-frame window comes back as 7 tokens; ask for 16 and you get 5 tokens, which actually covers 17 frames, because of the token grid. If that number surprises you, it's telling you why the seam you're staring at is a frame or two off.
Caching behaves the way you want
Both this node and the exact-frame loader declare IS_CHANGED as a SHA-256 digest of the file they read. Re-render the predecessor, the digest changes, and everything downstream re-executes on its own - no manual cache busting, and unlike the NaN trick a lot of loaders use, it doesn't poison the cache for the rest of the graph.
Install
Manager → Install Custom Nodes → search vrgamedev; or:
cd ComfyUI/custom_nodes
git clone https://github.com/vrgamegirl19/comfyui-vrgamedevgirl.git
python -m pip install -r comfyui-vrgamedevgirl/requirements.txt
Then restart ComfyUI and hard-refresh the browser tab. Fair warning: the requirements target the whole Video Builder - voxcpm, llama-cpp-python, demucs, av - so a fresh install pulls a lot of weight for a few hundred lines of latent slicing. On Windows portable with Python 3.13, run the README's bootstrap (upgrade pip/setuptools/wheel, then Cython scikit-build-core) or llama-cpp-python compiles from source. These nodes need a ComfyUI with comfy.nested_tensor; if that import fails you get a video-only latent, which joint audio-video sampling dislikes.
Where people get burned
Two things. First, a mismatch between context size here and the trim later - load 22 frames of context, trim 16 at the end, and you leave duplicated frames that read as a stutter. The Builder generates both numbers from one timing plan; hand-wiring, keep them identical.
Second, the frame/token grid. exact_frame_mode exists because the grid doesn't line up with what you see in the viewer. If your continuation starts a hair before or after the seam, that's why.
And the obvious one: predecessor latents only exist if you rendered with VRGDG H3 Save Latent in the graph. Add it to scene 1 before debugging scene 2.
Inputs (4)
| Name | Type | Default | Description |
|---|---|---|---|
| project_folder | STRING | Project folder path | |
| scene_number | INT | 11–999999 | Scene number to load (e.g. predecessor scene N-1) |
| context_frames | INT | 221–141 | Number of temporal context frames to slice from the tail (16, 22, 39, 56) |
| exact_frame_modeopt | BOOLEAN | false | Cut the context so it ends just before the predecessor's real last visible frame (skipping any padding after it) on the model's 5-token grid, leaving that last frame to be supplied as an exact image. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| context_latent | LATENT | — |
| sliced_tokens | INT | — |