Nodes/VRGameDevGirl Video Enhancement Nodes/VRGDG H3 Trim Continuation Output
ComfyUI Node

VRGDG H3 Trim Continuation Output

Cut the Duplicate Frames Off the Front of Every Continuation

By vrgamegirl19·Created about a year ago·Updated a day ago· 742
VRGDG H3 Trim Continuation Output
  • images
  • trimmed_images
◄context_frames22►

There's a price for continuity: when you feed scene 2 the last 22 frames of scene 1 as temporal context, H3 reproduces those frames at the start of scene 2's render. That's the cost of the handoff, and VRGDG H3 Trim Continuation Output is the receipt. It slices them off before the clip reaches your save node, so the scene you actually keep starts where the new content starts.

It's the least glamorous node in this part of the pack and arguably the most necessary one. Continuation without a trim isn't a feature, it's a stutter at every cut.

What it does, and what it deliberately doesn't

One input that carries behaviour: images, the decoded IMAGE batch coming out of your VAE Decode. context_frames (default 22) says how many leading frames to drop. The output is trimmed_images.

Under the hood it's a slice - images[context_frames:], cloned. No metadata lookups, no latent math, no cleverness. That's the honest design: the decoded batch is the wrong place to be clever. If the batch has fewer frames than you asked to cut, it prints a warning and hands you the whole thing unchanged, which is the right failure mode - a mis-set widget costs you a duplicated intro, not your render.

context_frames defaults to the same 22 the loader defaults to, and it accepts 0, which is a useful no-op when you're A/B-ing a graph with and without continuation.

Pick the number from the source, not the vibes

This is the one place people quietly get it wrong. The number you trim has to be the number of frames the context actually occupies, and that is not always the number you typed into the loader.

H3's temporal latent packs frames in token-sized bites (the first token carries about 1 frame, subsequent ones 4 each), so a "16 frame" context is really 17 frames across 5 tokens - the pack's own test suite asserts exactly that. Load 16, trim 16, and you keep one duplicated frame at the head of your scene. Load 22, trim 22, and you're clean, because 22 lines up. When in doubt, check the token count the loader reports and leave yourself a note about the frame count it maps to.

The bit the Builder does instead

In the pack's AI Video Builder, this node isn't in the generated graph at all - and that's not an oversight. The builder treats the context block as a warm-up in its timing plan and trims it from the video and the audio together, by the same offset, after the render. That matters because H3 generates picture and sound jointly: if the video loses 22 frames at the front and the audio doesn't, your lip sync is off by nearly a second and you'll spend an evening blaming the prompt.

So if you're using this node in a hand-built graph, be aware of what it can't do. It only knows about images. Trim the images, then trim the audio by the same 22 frames at 24 fps (about 0.92s) somewhere downstream - or better, do what the builder does and cut both in one timing step. Video-only trimming is the single most common way a perfectly good continuation sounds broken.

It also only helps the front. If your render carries cool-down frames after the scene's real end, this node won't touch them; that's a different trim, at the back.

Install

Manager → Install Custom Nodes → search vrgamedev, or:

cd ComfyUI/custom_nodes
git clone https://github.com/vrgamegirl19/comfyui-vrgamedevgirl.git
python -m pip install -r comfyui-vrgamedevgirl/requirements.txt

Restart ComfyUI, then hard-refresh the browser tab - the README calls that out because the pack's JavaScript UI won't reload otherwise. The pack's requirements.txt is heavy (it pulls voxcpm, llama-cpp-python, demucs, av and friends for the Builder and audio tools), so expect a slow first install even though this node needs nothing beyond torch. Windows portable users on Python 3.13: run the README's bootstrap first (pip install --upgrade pip setuptools wheel, then Cython scikit-build-core) or llama-cpp-python will attempt a source build.

Quick wiring

Sampler → VAE Decode → Video Combine / Save → (audio trim at the same offset)
                       └→ VRGDG H3 Trim Continuation Output → Save Video

Keep the untrimmed batch out of your final save path. It's very easy to leave two branches on the canvas, preview the wrong one, and conclude the trim node does nothing.

CategoryVRGDG/MiniMax H3 Latent Continuation

Inputs (2)

NameTypeDefaultDescription
imagesIMAGE—
context_framesINT220–999999Number of leading context frames to trim from the decoded output

Outputs (1)

NameTypeDescription
trimmed_imagesIMAGE—