ComfyUI Node

VRGDG H3 Save Latent

The Node You Slot Between Your Sampler and VAE Decode

By vrgamegirl19·Created about a year ago·Updated a day ago· 742
VRGDG H3 Save Latent
  • latent
  • latent
  • saved_path
◄project_folder►
◄scene_number1►
◄frame_count0►
◄fps24.0►
◄tail_padding_frames-1►

MiniMax H3 will give you 4–15 seconds per render. If you want a three-minute music video - which is exactly what people build with this pack - you don't ask it for three minutes. You render scene after scene and join them. How you join them is the whole argument. The pixel way is: decode scene 1, grab its last frame, feed that image back in as scene 2's start. The latent way is: never decode at all for the handoff, and give scene 2 the actual tensor scene 1 ended on as native temporal context.

That second path is what VRGDG H3 Save Latent exists for. It's the "write it down" half of latent continuation.

What it actually does

Feed it your sampled latent and it serializes it to <project_folder>/latents/scene_NNN.latent. Under the hood it pulls the video (and audio, when the render has it) out of the ComfyUI latent structure, moves it to CPU, and writes it with safetensors - falling back to torch.save if safetensors isn't importable. Alongside the tensors it records the stats a continuation needs later - frame count, token count, fps, video shape, whether audio was present, and tail_padding_frames if you set one - plus a small .json sidecar so a UI can read those numbers without loading the tensor.

The output is a passthrough. That's the design decision that matters: drop it between your sampler and VAE Decode and nothing downstream changes.

SamplerCustomAdvanced → VRGDG H3 Save Latent → VAE Decode → ...
                                ↓ latent
                        (next scene's Load Latent)

Why bother instead of round-tripping through pixels? A VAE decode → encode cycle is a lossy step you don't have to pay: same autoencoder, same latent space, so the tensor the sampler produced is already what the next render wants. Pixel handoffs also inherit whatever the decoder hallucinated in that last frame, and that compounds scene to scene.

The inputs that matter

Three are required. latent is your sampler output - it has to be the raw 5D H3 video latent [B, C, T, H, W], not a decoded image. project_folder is your project root, e.g. ComfyUI/output/MyVideo. scene_number is the 1-based index, and it decides the filename (scene_001.latent).

Then the optional three. frame_count defaults to 0, which means "auto-compute from the latent's token count" - leave it alone unless you know better. fps defaults to 24, which is H3's rate. tail_padding_frames defaults to -1, meaning unknown: it's how many frames at the end of this render fall after the scene's real visible last frame (alignment padding or a cool-down tail). If you plan to use exact-frame continuation later, whoever renders the predecessor has to record that number, or the next scene can't find the true last frame inside the latent.

Outputs are latent (the passthrough, straight into VAE Decode) and saved_path (the absolute path as a string - handy for a note node, or for feeding a manual load step on another machine).

Heads-up, this is where people get burned

It does not scream when it fails. If project_folder is empty or wrong, or the latent isn't a 5D tensor, you get a console warning and saved_path comes back empty - the run still reports success. If your next scene later errors with a missing latent, check this warning line first.

Don't feed it a 4D image latent. An image latent or an already-decoded tensor fails the shape check and no file is written. For an image-based handoff this is the wrong node.

Note the frame–token arithmetic. H3 frames arrive in token-sized bites (the pack's table cycles 1 frame then 4 per token), so 22 frames is 7 tokens and a "16 frame" context is really 17 frames. The sidecar JSON tells you what it computed.

Install

ComfyUI Manager → Install Custom Nodes → search vrgamedev, or clone it by hand:

cd ComfyUI/custom_nodes
git clone https://github.com/vrgamegirl19/comfyui-vrgamedevgirl.git
python -m pip install -r comfyui-vrgamedevgirl/requirements.txt

Restart ComfyUI and hard-refresh the browser. The pack's requirements.txt is bigger than these five nodes need - it drags in voxcpm, llama-cpp-python, demucs and av for the Video Builder and its audio tools - and on Windows portable with Python 3.13 the README's bootstrap (upgrade pip/setuptools/wheel, then Cython scikit-build-core) is the difference between installing cleanly and watching llama-cpp-python compile itself. These continuation nodes only need torch, safetensors, numpy and PIL, plus a ComfyUI new enough to have comfy.nested_tensor.

Where it sits in a real graph

Scene 1 renders normally - Save Latent sits before VAE Decode and drops scene_001.latent on disk. Scene 2 loads that file, slices its tail as context, and injects it as a keyframe. You never touched a PNG. Render in order, though: scene 2 can't continue from a scene 1 that hasn't been rendered, and re-rendering scene 1 makes scene 2's continuation stale by definition.

CategoryVRGDG/MiniMax H3 Latent Continuation

Inputs (6)

NameTypeDefaultDescription
latentLATENT—
project_folderSTRINGProject folder path (e.g. ComfyUI/output/MyProject)
scene_numberINT11–999999Scene number (1-based index)
frame_countoptINT00–999999Optional explicit frame count (0 = auto-compute from latent tokens)
fpsoptFLOAT24.01–120—
tail_padding_framesoptINT-1-1–999999Frames at the end of this render that fall after the scene's visible last frame (alignment padding / cool-down). -1 = unknown. Lets the next scene's exact-frame continuation find the real last frame inside the latent.

Outputs (2)

NameTypeDescription
latentLATENT—
saved_pathSTRING—